Назад
Company hidden
2 дня назад

Machine Learning Operations (MLOps) Engineer

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Operations (MLOps) Engineer (Python/Cloud/Kubernetes): Building and operating automated ML pipelines and platform infrastructure that move models from experimentation into reliable, cost-efficient production with an accent on CI/CD, observability, reproducibility, and safe deployment. Focus on designing scalable training and inference workflows, implementing canary and rollback strategies, diagnosing model degradation, and reducing operational toil through reusable automation.

Location: Remote employee; role listed for the United States, US - IN - Carmel

Company

hirify.global operates a digital marketplace for wholesale used vehicles and provides data-driven tools that help customers buy and sell more effectively.

What you will do

  • Design, build, and operate automated ML pipelines for data ingestion, training, validation, deployment, and rollback.
  • Develop model CI/CD workflows with automated testing, canary releases, and rollback strategies.
  • Implement monitoring and observability for model performance, data drift, inference latency, and service availability.
  • Ensure reproducibility through versioned code, data, features, and execution environments.
  • Automate operational work and create reusable ML platform components, documentation, tests, and runbooks.
  • Translate data science and product needs into scalable ML delivery designs while partnering with infrastructure and engineering teams.

Requirements

  • 3+ years of experience in MLOps, ML engineering, SRE, DevOps, or software engineering, including production ML pipelines.
  • Strong Python skills and hands-on experience with a major cloud platform, preferably AWS.
  • Experience with workflow orchestration, infrastructure as code, and CI/CD systems.
  • Experience with containers and Kubernetes, with Docker and EKS preferred.
  • Familiarity with model-serving frameworks and monitoring tools such as TorchServe, Triton, BentoML, SageMaker endpoints, Prometheus, Grafana, or Evidently AI.
  • Production experience with system debugging, observability, incident response, and critical evaluation of AI-assisted development output.

Nice to have

  • Experience with SageMaker, MLflow, Databricks, Vertex AI, feature stores, or model registries.
  • Experience deploying and evaluating LLM or generative AI systems.
  • Experience with GPU workload scheduling and inference cost optimization.
  • Experience interviewing candidates, conducting technical debriefs, or mentoring junior engineers and interns.

Culture & Benefits

  • Remote work with occasional team alignment meetings.
  • Medical, dental, and vision benefits with US HSA contributions and FSA options.
  • Immediately vested 401(k) with company match in the US.
  • Paid vacation, personal, sick, maternity, and paternity leave, with additional disability and life insurance benefits in the US.
  • Tuition reimbursement, employee assistance resources, and opportunities for knowledge sharing and career advancement.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →