4 дня назад
GPU Cluster Infrastructure Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Cluster Infrastructure Engineer (AWS HyperPod): Operating and evolving large-scale AWS GPU clusters for foundation model training and customer fine-tuning workloads with an accent on cluster reliability, orchestration, and GPU utilization. Focus on automated job recovery, distributed training workflows, self-service tooling, and cost and capacity optimization across Spot Instances, Reserved Instances, and Savings Plans.
Location: On-site in Metzingen / Riederich, Germany
Company
develops robotics products supported by advanced software engineering and machine learning infrastructure.
What you will do
- Operate and continuously evolve large-scale AWS HyperPod GPU clusters using Slurm and EKS orchestration models.
- Design cluster-stability mechanisms, including node-failure detection, automated job recovery, checkpoint coordination, and fault-tolerant multi-node training.
- Optimize GPU utilization across compute, memory, EFA networking, and storage throughput.
- Build self-service tooling for ML researchers and engineers to launch, monitor, and manage training workloads.
- Define workload-priority, capacity, and cost-management strategies across pretraining, fine-tuning, and customer workloads.
- Collaborate with AWS HyperPod teams and internal ML, product, finance, and cloud vendor stakeholders.
Requirements
- 5+ years of infrastructure or systems engineering experience, with a strong focus on GPU cluster or HPC operations.
- Hands-on experience with AWS HyperPod and AWS GPU instances.
- Strong knowledge of Slurm and Kubernetes and the trade-offs between them for large-scale GPU workloads.
- Practical understanding of distributed training, including throughput analysis and debugging.
- Experience building self-service tooling and operational documentation for technical users.
- Strong English communication skills required; German is a plus.
Culture & Benefits
- Infrastructure changes and cluster configurations are managed through Infrastructure as Code.
- Work directly with AWS product and solutions engineering teams to report issues and influence the HyperPod roadmap.
- Create onboarding documentation, training materials, and internal workshops for ML infrastructure users.
- Manage Spot interruption handling, capacity reservations, cost attribution, Reserved Instances, Savings Plans, and AWS commitment negotiations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Machine Learning & Cloud Infra Engineer (AI)
4 дня назад
Senior DevOps Engineer (Kubernetes)
5 дней назад
Private Cloud Engineer (Kubernetes)
5 дней назад
Cloud Platform Engineer (Fintech)
4 дня назад
Senior Cloud Site Reliability Engineer (Platform) (m/f/x)
4 дня назад