5 дней назад
Staff ML Platform Engineer (AWS/Databricks)
224 000 - 280 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff ML Platform Engineer (AWS/Databricks): Building and operating the paved-road ML platform for training, inference, LLM serving, and observability across regulated healthcare data with an accent on Databricks, SageMaker, MLflow, Spark, and PHI-safe infrastructure. Focus on architecting self-service production workflows, solving GPU and distributed-compute challenges, and designing reliable, cost-aware LLM serving across managed and self-hosted environments.
Location: Remote - United States
Salary: $224,000–$280,000 USD per year
Company
is a healthcare data collaboration platform that provides secure data solutions for providers, health plans, researchers, and life sciences organizations.
What you will do
- Set technical direction for ML training, serving, observability, GPU capacity, Spark tuning, and production incident resolution.
- Evolve the shared CI/CD spine, model-workflow scaffolding, and Databricks Asset Bundles into a self-service ML platform.
- Lead architecture for LLM endpoint serving across Databricks, AWS, Snowflake, and self-hosted deployments, including latency, cost, caching, evaluation, and PHI-safe routing.
- Define standards and tooling for MLflow, model registries, training image supply chains, and training and inference observability.
- Partner with Data Science, Application Development, Operations, vendors, and platform teams on architecture and technology selection.
- Mentor senior engineers while remaining hands-on with production code and Infrastructure-as-Code.
Requirements
- 10+ years of software engineering experience, including 3+ years designing and operating enterprise-scale ML platforms in production.
- Production experience with Databricks and/or Amazon SageMaker, MLflow or an equivalent system, and a core ML framework such as PyTorch or TensorFlow.
- Strong Java or JVM-equivalent and Python skills, with deep Apache Spark experience in large-scale distributed computing.
- Deep AWS experience across networking, IAM, GPU compute, storage, and messaging services.
- Fluency with Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads.
- Production LLM serving experience, including cost management, evaluation, and safe handling of sensitive prompts and outputs.
Nice to have
- Technical leadership on healthcare or regulated-industry ML platforms involving HIPAA, HITRUST, SOC 2, or equivalent standards.
- Experience with Databricks Asset Bundles, Unity Catalog, Iceberg, Delta, Kafka, or Kinesis.
- Experience with specialized inference pipelines, GPU capacity planning, real-time inference, safety-sensitive AI evaluation, or open-source ML infrastructure.
Culture & Benefits
- Collaborative, high-performance environment focused on transforming healthcare through data logistics products and services.
- Remote work in the United States.
- Full-time employment with total rewards compensation.
- Post-offer health screenings and vaccination requirements may apply depending on client and state requirements.
- Reasonable accommodations are available for candidates with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Lead MLOps Engineer (Databricks)
100 000 - 200 000$
9 дней назад
Senior Machine Learning Systems Engineer (CAD)
154 000 - 193 000CAD
9 дней назад
Sr. ML Engineer (AI)
123 400 - 191 100$
9 дней назад
Staff AI Engineer
205 000 - 307 000$
6 дней назад
ML Infrastructure Engineer (AI)
6 дней назад
Software Development Engineer - ML Ops (US Federal)
163 800 - 245 800$