4 дня назад
ML Infrastructure Engineer (AI)
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Infrastructure Engineer (AI): Designing and operating GPU infrastructure and platform systems for large-scale AI training and inference with an accent on distributed training, scheduling, storage performance, and reliability. Focus on integrating ML frameworks, building high-performance networking and observability, optimizing infrastructure costs, and implementing fault tolerance for multi-tenant workloads.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate GPU and accelerator infrastructure across on-premises clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, resource-sharing, and automation systems to improve accelerator utilization across teams.
- Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, Ray Train, and related ML training frameworks.
- Operate high-performance storage, data pipelines, RDMA/InfiniBand networking, NCCL, and collective communication systems.
- Build observability, checkpointing, restart, fault-tolerance, security, and isolation capabilities for multi-tenant AI workloads.
- Develop researcher tooling, capacity planning processes, operational documentation, and cost-optimization strategies.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 6+ years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong Python skills and proficiency in at least one systems language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, collective communication, Linux internals, networking, and high-performance storage.
- Experience with Kubernetes, Slurm, Ray, a major cloud provider’s ML infrastructure, testing, CI/CD, and code review.
Nice to have
- Experience operating InfiniBand or RDMA networking at scale.
- Open-source ML infrastructure contributions or familiarity with custom orchestrators and research-grade training stacks.
- Exposure to frontier model training operations and FinOps for AI workloads.
Culture & Benefits
- Full-time direct W-2 employment.
- Career growth opportunity within an established organization.
- Remote work arrangement within the United States.
- New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Benzinga
2 дня назад
AI Engineer (Go)
90 000 - 100 000$
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
Airbnb
6 дней назад
Senior ML Engineer (AI Engineering)
191 000 - 223 000$
6 дней назад
Staff Software Engineer (Imaging ML)
219 000 - 233 000$
5 дней назад
AI Infrastructure Engineer
170 500 - 315 490$
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$