2 часа назад
Machine Learning & Cloud Infra Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning & Cloud Infra Engineer (AI): Building and operating scalable infrastructure for large diffusion-based generative models with an accent on GPU clusters, distributed training, storage, orchestration, and reliable model serving. Focus on optimizing petabyte-scale data throughput, enabling PyTorch distributed training, and designing secure, observable systems for research-to-production workflows.
Location: On-site in London or Munich
Company
develops World Models that combine generative AI, computer vision, and simulation to create physically grounded 3D environments.
What you will do
- Own and evolve ML and cloud infrastructure for training and evaluating large foundation models.
- Provision, scale, and maintain multi-node, multi-GPU clusters across cloud and on-premises environments.
- Enable high-throughput distributed training with PyTorch DDP/FSDP, NCCL, and performance debugging.
- Build storage, networking, caching, and data-locality systems for petabyte-scale datasets.
- Deploy workloads with Docker and Kubernetes, maintain Terraform infrastructure, and support reliable release processes.
- Implement observability, security, access management, incident response, and model-serving pathways.
Requirements
- 3+ years of professional experience in infrastructure, platform, or cloud engineering.
- Hands-on experience with GPU compute, CUDA/NCCL concepts, utilization analysis, profiling, and networking bottlenecks.
- Strong experience with AWS, GCP, or Azure, including networking, IAM, and cost management.
- Proficiency with Docker, Kubernetes, Terraform, Python, and Bash or PowerShell.
- Familiarity with PyTorch and distributed training approaches such as DDP/FSDP.
- Experience with monitoring, observability, and CI/CD tooling for infrastructure or ML workflows.
Culture & Benefits
- Work in a small engineering team focused on generative 3D AI and frontier systems.
- Collaborate closely with ML researchers and engineers on developer tooling and training infrastructure.
- Contribute to systems spanning robotics, AR/VR, gaming, and cinema applications.
- Inclusive workplace committed to equal opportunity and fair treatment throughout recruitment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 часов назад
Infrastructure Engineer (AI)
13 часов назад
Infrastructure Engineer (AI)
13 часов назад
Platform Engineer (AI)
13 часов назад
Software Engineer (DevOps/Platform)
12 часов назад
Senior DevOps Engineer (Cloud)
3 дня назад
Software Engineering, Machine Learning Operations, Tapestry (AI)
166 000 - 244 000$