4 дня назад
Machine Learning Systems & Infrastructure Engineer (Generative AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning Systems & Infrastructure Engineer (Generative AI): Building scalable training stacks, data ingestion pipelines, experiment orchestration, and production endpoints for large diffusion-based world models with an accent on distributed GPU training, petabyte-scale storage, and reliable ML infrastructure. Focus on optimizing PyTorch DDP/FSDP workloads, designing end-to-end dataset pipelines, and operating Kubernetes-, Terraform-, and observability-based systems.
Location: On-site in London or Munich
Company
develops generative AI, computer vision, and simulation systems for physically grounded 3D world models used across robotics, AR/VR, gaming, and cinema.
What you will do
- Build and operate ML systems for training, evaluation, checkpointing, and serving large foundation models.
- Improve distributed PyTorch training stacks using DDP/FSDP and NCCL for performance, stability, reproducibility, and preemption-safe checkpointing.
- Develop Python data ingestion and preprocessing pipelines that create clean, versioned datasets from third-party capture sources.
- Operate workflow orchestration, GPU scheduling, experiment tracking, managed inference, and launcher SDK systems.
- Ship containerized workloads with Docker and Kubernetes, maintain Terraform infrastructure and CI/CD pipelines, and support self-hosted GPU runners.
- Establish monitoring, logging, alerting, SLOs, incident response, security, IAM, and network boundaries for owned systems.
Requirements
- 3+ years of production-quality Python development in a large, multi-author codebase.
- Hands-on experience with PyTorch and distributed training using DDP/FSDP or comparable technologies, including debugging jobs across multiple GPUs and nodes.
- Experience shipping end-to-end data pipelines covering ingestion, transformation, validation, versioning, and republishing at scale.
- Experience with GPU performance debugging, CUDA/NCCL, networking bottlenecks, profiling, and cloud environments such as AWS, GCP, or Azure.
- Proficiency with Docker and Kubernetes, plus the ability to write and maintain Terraform infrastructure and CI/CD workflows.
- Knowledge of SQL, relational, analytical, embedded, and object-storage systems, with familiarity with ML orchestration and experiment tracking.
Culture & Benefits
- Work in a small engineering and research team focused on difficult generative 3D AI problems.
- Collaborate closely with ML researchers, engineers, and the platform team in a shared monorepo.
- Contribute to an inclusive and equal-opportunity workplace welcoming people from all backgrounds and perspectives.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →