Назад
Company hidden
4 дня назад

GPU Cluster Infrastructure Engineer

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Cluster Infrastructure Engineer (AWS HyperPod): Operating and evolving large-scale AWS GPU clusters for foundation model training and customer fine-tuning workloads with an accent on cluster reliability, orchestration, and GPU utilization. Focus on automated job recovery, distributed training workflows, self-service tooling, and cost and capacity optimization across Spot Instances, Reserved Instances, and Savings Plans.

Location: On-site in Metzingen / Riederich, Germany

Company

hirify.global develops robotics products supported by advanced software engineering and machine learning infrastructure.

What you will do

  • Operate and continuously evolve large-scale AWS HyperPod GPU clusters using Slurm and EKS orchestration models.
  • Design cluster-stability mechanisms, including node-failure detection, automated job recovery, checkpoint coordination, and fault-tolerant multi-node training.
  • Optimize GPU utilization across compute, memory, EFA networking, and storage throughput.
  • Build self-service tooling for ML researchers and engineers to launch, monitor, and manage training workloads.
  • Define workload-priority, capacity, and cost-management strategies across pretraining, fine-tuning, and customer workloads.
  • Collaborate with AWS HyperPod teams and internal ML, product, finance, and cloud vendor stakeholders.

Requirements

  • 5+ years of infrastructure or systems engineering experience, with a strong focus on GPU cluster or HPC operations.
  • Hands-on experience with AWS HyperPod and AWS GPU instances.
  • Strong knowledge of Slurm and Kubernetes and the trade-offs between them for large-scale GPU workloads.
  • Practical understanding of distributed training, including throughput analysis and debugging.
  • Experience building self-service tooling and operational documentation for technical users.
  • Strong English communication skills required; German is a plus.

Culture & Benefits

  • Infrastructure changes and cluster configurations are managed through Infrastructure as Code.
  • Work directly with AWS product and solutions engineering teams to report issues and influence the HyperPod roadmap.
  • Create onboarding documentation, training materials, and internal workshops for ML infrastructure users.
  • Manage Spot interruption handling, capacity reservations, cost attribution, Reserved Instances, Savings Plans, and AWS commitment negotiations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →