4 дня назад
AI Infrastructure Engineer
100 000 - 160 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Infrastructure Engineer (GPU/ML Infrastructure): Designing and operating the platform layer for large-scale AI training and inference with an accent on GPU clusters, distributed training, scheduling, storage performance, and reliability. Focus on building accelerator resource-sharing systems, integrating ML frameworks, optimizing high-performance networking and storage, and implementing fault-tolerant workloads at scale.
Location: 100% remote within the United States
Salary: $100,000–$160,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design and operate GPU and accelerator infrastructure across on-premises clusters, cloud-managed services, and hybrid configurations.
- Build scheduling, queueing, resource-sharing, and developer tooling systems for large-scale ML workloads.
- Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified AI platform.
- Operate high-performance storage, data pipelines, RDMA, InfiniBand, NCCL, and high-bandwidth networking architectures.
- Implement observability, checkpointing, restart, fault tolerance, security controls, isolation, and access management for multi-tenant infrastructure.
- Drive automation, capacity planning, operational documentation, and cost optimization across compute, storage, and networking.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 10+ years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong proficiency in Python and at least one systems language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, collective communication, Linux internals, networking, and high-performance storage.
- Must be authorized to work in the United States; eligible profiles include U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates. New H-1B visa petitions cannot be sponsored.
Nice to have
- Experience operating InfiniBand or RDMA networking at scale.
- Contributions to open-source ML infrastructure projects.
- Experience with custom orchestrators, research-grade training stacks, frontier model training operations, or FinOps for AI workloads.
Culture & Benefits
- Full-time direct W-2 employment.
- Remote work within the United States.
- Career growth opportunities within an established organization.
- Equal employment opportunity and a workplace free from harassment and discrimination.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
Benzinga
2 дня назад
AI Engineer (Go)
90 000 - 100 000$
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
4 дня назад
AI Software Engineer
120 000 - 150 000$
2 дня назад
Senior AI Engineer (Applied AI Solutions)
195 000 - 225 000$
2 дня назад
Applied AI Engineer
125 000 - 155 000$