2 дня назад
Infrastructure, Large-scale Training (AI)
180 000 - 450 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure, Large-scale Training (AI) (GPU Clusters and ML Infrastructure): Building and operating large-scale GPU computing clusters for AI training and inference with an accent on reliability, scalability, and cost efficiency. Focus on provisioning 10,000+ GPU environments, optimizing networking and scheduling, and solving complex distributed-systems and production-reliability challenges.
Location: San Jose, United States
Salary: $180,000–$450,000 annually
Company
is an artificial intelligence company developing proactive, multimodal intelligence and next-generation hardware that interact with people and the real world through speech, text, vision, and persistent memory.
What you will do
- Design and maintain Infrastructure as Code practices for repeatable, auditable, and scalable GPU cluster provisioning.
- Improve and harden CI/CD pipelines for secure, reliable, low-latency model delivery in production.
- Own training infrastructure operating at the scale of 10,000+ GPUs, including scheduling, fault tolerance, and network fabric optimization.
- Partner with ML researchers and engineers to identify compute bottlenecks and deliver infrastructure improvements.
- Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.
- Drive capacity planning, cost efficiency, hardware lifecycle management, and internal tooling for compute users.
Requirements
- 5+ years of experience in infrastructure, systems, or platform engineering, including at least 2 years in ML or HPC environments.
- Experience managing GPU clusters or large-scale distributed compute infrastructure.
- Strong proficiency in at least one systems or infrastructure programming language.
- Deep understanding of networking fundamentals relevant to high-throughput training workloads.
- Experience with container orchestration, job scheduling, and multi-tenant resource management.
- Production systems ownership, high-reliability operations, debugging, and observability across the infrastructure stack.
Nice to have
- Experience operating large, GPU-aware Kubernetes clusters.
- Pulumi or similar Infrastructure as Code tooling.
- Rust or Go for systems-level tooling and performance-critical services.
- Familiarity with PyTorch and Ray.
- RDMA, InfiniBand, or RoCE experience.
Culture & Benefits
- Full-time position with a US annual base salary range of $180,000–$450,000.
- Total compensation may include additional components and benefits depending on the role.
- Work centers on highly technical infrastructure-as-a-product engineering for AI research and production workloads.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Foundational Software Engineer (Infrastructure)
180 000 - 300 000$
2 дня назад
Supercomputing Platform & Infrastructure Engineer (AI)
200 000 - 550 000$
2 дня назад
Infrastructure Engineer (AI)
150 000 - 300 000CAD
2 дня назад
Infrastructure Engineer (AI)
148 000 - 230 000$
2 дня назад
Infrastructure Engineer (AI Security)
180 000 - 290 000$
2 дня назад
Infrastructure Engineer (AI)
200 000 - 400 000$