2 дня назад
Machine Learning Operations (MLOps) Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning Operations (MLOps) Engineer (Python/Cloud/Kubernetes): Building and operating automated ML pipelines and platform infrastructure that move models from experimentation into reliable, cost-efficient production with an accent on CI/CD, observability, reproducibility, and safe deployment. Focus on designing scalable training and inference workflows, implementing canary and rollback strategies, diagnosing model degradation, and reducing operational toil through reusable automation.
Location: Remote employee; role listed for the United States, US - IN - Carmel
Company
operates a digital marketplace for wholesale used vehicles and provides data-driven tools that help customers buy and sell more effectively.
What you will do
- Design, build, and operate automated ML pipelines for data ingestion, training, validation, deployment, and rollback.
- Develop model CI/CD workflows with automated testing, canary releases, and rollback strategies.
- Implement monitoring and observability for model performance, data drift, inference latency, and service availability.
- Ensure reproducibility through versioned code, data, features, and execution environments.
- Automate operational work and create reusable ML platform components, documentation, tests, and runbooks.
- Translate data science and product needs into scalable ML delivery designs while partnering with infrastructure and engineering teams.
Requirements
- 3+ years of experience in MLOps, ML engineering, SRE, DevOps, or software engineering, including production ML pipelines.
- Strong Python skills and hands-on experience with a major cloud platform, preferably AWS.
- Experience with workflow orchestration, infrastructure as code, and CI/CD systems.
- Experience with containers and Kubernetes, with Docker and EKS preferred.
- Familiarity with model-serving frameworks and monitoring tools such as TorchServe, Triton, BentoML, SageMaker endpoints, Prometheus, Grafana, or Evidently AI.
- Production experience with system debugging, observability, incident response, and critical evaluation of AI-assisted development output.
Nice to have
- Experience with SageMaker, MLflow, Databricks, Vertex AI, feature stores, or model registries.
- Experience deploying and evaluating LLM or generative AI systems.
- Experience with GPU workload scheduling and inference cost optimization.
- Experience interviewing candidates, conducting technical debriefs, or mentoring junior engineers and interns.
Culture & Benefits
- Remote work with occasional team alignment meetings.
- Medical, dental, and vision benefits with US HSA contributions and FSA options.
- Immediately vested 401(k) with company match in the US.
- Paid vacation, personal, sick, maternity, and paternity leave, with additional disability and life insurance benefits in the US.
- Tuition reimbursement, employee assistance resources, and opportunities for knowledge sharing and career advancement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
AWS Cloud DevOps Engineer (AI/ML)
5 дней назад
Senior MLOps Engineer (Fintech)
1 100 - 1 600PLN
3 дня назад
Director, MLOps
210 000 - 330 000$
3 дня назад
Senior AI/ML Software Engineer
65 600 - 98 400€
СИНЕРГИЯ
5 дней назад
DevOps инженер / MLOps (AI)
300 000₽
7 дней назад