Назад
Company hidden
4 дня назад

Machine Learning Systems & Infrastructure Engineer (Generative AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
UK/Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Systems & Infrastructure Engineer (Generative AI): Building scalable training stacks, data ingestion pipelines, experiment orchestration, and production endpoints for large diffusion-based world models with an accent on distributed GPU training, petabyte-scale storage, and reliable ML infrastructure. Focus on optimizing PyTorch DDP/FSDP workloads, designing end-to-end dataset pipelines, and operating Kubernetes-, Terraform-, and observability-based systems.

Location: On-site in London or Munich

Company

hirify.global develops generative AI, computer vision, and simulation systems for physically grounded 3D world models used across robotics, AR/VR, gaming, and cinema.

What you will do

  • Build and operate ML systems for training, evaluation, checkpointing, and serving large foundation models.
  • Improve distributed PyTorch training stacks using DDP/FSDP and NCCL for performance, stability, reproducibility, and preemption-safe checkpointing.
  • Develop Python data ingestion and preprocessing pipelines that create clean, versioned datasets from third-party capture sources.
  • Operate workflow orchestration, GPU scheduling, experiment tracking, managed inference, and launcher SDK systems.
  • Ship containerized workloads with Docker and Kubernetes, maintain Terraform infrastructure and CI/CD pipelines, and support self-hosted GPU runners.
  • Establish monitoring, logging, alerting, SLOs, incident response, security, IAM, and network boundaries for owned systems.

Requirements

  • 3+ years of production-quality Python development in a large, multi-author codebase.
  • Hands-on experience with PyTorch and distributed training using DDP/FSDP or comparable technologies, including debugging jobs across multiple GPUs and nodes.
  • Experience shipping end-to-end data pipelines covering ingestion, transformation, validation, versioning, and republishing at scale.
  • Experience with GPU performance debugging, CUDA/NCCL, networking bottlenecks, profiling, and cloud environments such as AWS, GCP, or Azure.
  • Proficiency with Docker and Kubernetes, plus the ability to write and maintain Terraform infrastructure and CI/CD workflows.
  • Knowledge of SQL, relational, analytical, embedded, and object-storage systems, with familiarity with ML orchestration and experiment tracking.

Culture & Benefits

  • Work in a small engineering and research team focused on difficult generative 3D AI problems.
  • Collaborate closely with ML researchers, engineers, and the platform team in a shared monorepo.
  • Contribute to an inclusive and equal-opportunity workplace welcoming people from all backgrounds and perspectives.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →