Назад
Company hidden
обновлено 5 дней назад

Senior ML Infrastructure & MLOps Engineer (Core Platform)

126 900 - 185 100$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior ML Infrastructure & MLOps Engineer (Core Platform) (ML infrastructure/MLOps): Building shared systems for the machine learning lifecycle, including standardized training frameworks, feature-generation platforms, and high-performance model-serving clusters with an accent on GKE orchestration, distributed training, and automated ML pipelines. Focus on designing scalable infrastructure, optimizing model-training workflows, reducing cloud costs, and stabilizing complex production deployments.

Location: La Honda, California, United States

Salary: $126,900–$185,100 per year

Company

hirify.global provides staffing and technical workforce solutions.

What you will do

  • Architect and maintain high-performance ML training and model-serving infrastructure on Google Kubernetes Engine.
  • Build optimization pipelines, including knowledge distillation and foundational training tooling.
  • Develop automated pipelines for model training, validation, and continuous deployment.
  • Create scalable data-sampling and feature-generation platforms for ML research and experimentation.
  • Build standardized deployment tools that improve platform adoption and onboarding for research and engineering teams.
  • Collaborate with ML researchers and software engineers to turn theoretical models into scalable production systems.

Requirements

  • 5–10+ years of experience designing and operating large-scale distributed ML platforms.
  • Experience supporting production-grade ML workflows in cloud environments.
  • Deep expertise in container orchestration, especially GKE or equivalent enterprise Kubernetes environments.
  • Hands-on experience with scalable ML pipelines such as Kubeflow, Airflow, or TFX.
  • Strong proficiency in distributed training, feature stores, and model-serving infrastructure.
  • Ownership-driven, pragmatic approach with strong communication and collaboration skills.

Nice to have

  • Experience on tier-one enterprise ML/AI platform teams.
  • Knowledge of distributed-systems backend optimization and infrastructure as code.

Culture & Benefits

  • Focus on reliability, infrastructure uptime, cost efficiency, and developer velocity.
  • Cross-functional collaboration with researchers and platform engineers.
  • For temporary assignments lasting 13 weeks or longer: medical, dental, vision, and 401(k) benefits.
  • Statutory sick pay where required.
  • Reasonable accommodations are available during the employment process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →