Назад
Company hidden
5 дней назад

Staff ML Platform Engineer (AWS/Databricks)

224 000 - 280 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff ML Platform Engineer (AWS/Databricks): Building and operating the paved-road ML platform for training, inference, LLM serving, and observability across regulated healthcare data with an accent on Databricks, SageMaker, MLflow, Spark, and PHI-safe infrastructure. Focus on architecting self-service production workflows, solving GPU and distributed-compute challenges, and designing reliable, cost-aware LLM serving across managed and self-hosted environments.

Location: Remote - United States

Salary: $224,000–$280,000 USD per year

Company

hirify.global is a healthcare data collaboration platform that provides secure data solutions for providers, health plans, researchers, and life sciences organizations.

What you will do

  • Set technical direction for ML training, serving, observability, GPU capacity, Spark tuning, and production incident resolution.
  • Evolve the shared CI/CD spine, model-workflow scaffolding, and Databricks Asset Bundles into a self-service ML platform.
  • Lead architecture for LLM endpoint serving across Databricks, AWS, Snowflake, and self-hosted deployments, including latency, cost, caching, evaluation, and PHI-safe routing.
  • Define standards and tooling for MLflow, model registries, training image supply chains, and training and inference observability.
  • Partner with Data Science, Application Development, Operations, vendors, and platform teams on architecture and technology selection.
  • Mentor senior engineers while remaining hands-on with production code and Infrastructure-as-Code.

Requirements

  • 10+ years of software engineering experience, including 3+ years designing and operating enterprise-scale ML platforms in production.
  • Production experience with Databricks and/or Amazon SageMaker, MLflow or an equivalent system, and a core ML framework such as PyTorch or TensorFlow.
  • Strong Java or JVM-equivalent and Python skills, with deep Apache Spark experience in large-scale distributed computing.
  • Deep AWS experience across networking, IAM, GPU compute, storage, and messaging services.
  • Fluency with Terraform, containers, Kubernetes, and GitHub-based CI/CD for ML workloads.
  • Production LLM serving experience, including cost management, evaluation, and safe handling of sensitive prompts and outputs.

Nice to have

  • Technical leadership on healthcare or regulated-industry ML platforms involving HIPAA, HITRUST, SOC 2, or equivalent standards.
  • Experience with Databricks Asset Bundles, Unity Catalog, Iceberg, Delta, Kafka, or Kinesis.
  • Experience with specialized inference pipelines, GPU capacity planning, real-time inference, safety-sensitive AI evaluation, or open-source ML infrastructure.

Culture & Benefits

  • Collaborative, high-performance environment focused on transforming healthcare through data logistics products and services.
  • Remote work in the United States.
  • Full-time employment with total rewards compensation.
  • Post-offer health screenings and vaccination requirements may apply depending on client and state requirements.
  • Reasonable accommodations are available for candidates with disabilities.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →