Назад

Не получаете ответ?

Telegram-вакансии старше 7 дней могут быть уже неактуальны.

14 дней назад

Software Engineer ML Ops and Platform

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Страна
Germany
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Software Engineer ML Ops and Platform (Machine Learning): Building production ML infrastructure and self-service platforms for training, evaluation, telemetry, deployment, and governed model artifacts with an accent on multi-tenant GPU clusters, CI/CD, and secure data lineage. Focus on staged rollouts, rollback, drift detection, real-time monitoring, reproducibility, privacy, and edge delivery.

Software Engineer ML Ops and Platform

Company

Tools for Humanity

Conditions

5 days ago

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own the machine learning lifecycle from data to device. You will design and operate production pipelines for training evaluation telemetry and deployment, build self-service platforms, enable staged rollouts and rollback, expose governed datasets and model artifacts through secure services, and implement monitoring lineage reproducibility and privacy controls.

Requirements

  • 5+ years of experience building ML infrastructure data platforms or production ML systems at scale
  • Experience delivering platforms and CI/CD pipelines used by ML or data teams
  • Experience running large-scale training on multi-tenant GPU clusters
  • Experience building versioned dataset and lineage systems with slice-level provenance and governed access
  • Deep understanding of Docker Kubernetes or EKS and Infrastructure-as-Code tools such as Terraform CDK or CloudFormation
  • Strong backend engineering skills in Python and/or Go
  • Understanding of modern CI/CD model packaging and observability practices
  • Experience operating production systems defining SLAs and handling rollout or incident workflows
  • Comfort using modern agentic AI development

Responsibilities

  • Design and operate infrastructure for training evaluation telemetry ingestion and deployment
  • Maintain CI/CD workflows and automated pipelines
  • Build edge-aware rollout services with staged deployment A/B experimentation and rollback
  • Develop secure APIs and backend services for governed datasets and model artifacts
  • Implement automated checks drift detection and real-time model monitoring
  • Establish practices for data lineage reproducibility privacy security and secure edge delivery
  • Collaborate with ML research product and firmware teams to improve delivery workflows

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -