3 дня назад
Intelligent Infrastructure Engineer (AI)
100 000 - 150 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Intelligent Infrastructure Engineer (AI): Designing, building, and operating the platform layer for large-scale AI training and inference workloads with an accent on GPU clusters, distributed training, scheduling, storage performance, and developer experience. Focus on connecting hardware, kernels, schedulers, and ML frameworks while improving reliability, efficiency, and cost control.
Location: 100% remote within the United States
Salary: $100,000–$150,000 annually
Company
is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
What you will do
- Design, build, and operate platform infrastructure for large-scale AI training and inference workloads.
- Operate and optimize GPU clusters, distributed training frameworks, scheduling systems, and high-performance storage.
- Improve infrastructure reliability, efficiency, developer experience, and cost control for ML engineers and researchers.
- Work across hardware, kernels, schedulers, ML frameworks, networking, and cloud infrastructure.
- Apply software engineering practices including testing, CI/CD, and code review.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- 6+ years of experience in infrastructure, platform, or HPC engineering.
- Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
- Strong Python skills and proficiency in at least one systems language, such as Go or C++.
- Experience with distributed training, accelerator architectures, collective communication, Kubernetes, Slurm, Ray, or similar ML scheduling systems.
- Strong knowledge of Linux internals, networking, high-performance storage, and at least one major cloud provider’s ML infrastructure offerings.
Nice to have
- Experience operating InfiniBand or RDMA networking at scale.
- Contributions to open-source ML infrastructure projects.
- Familiarity with custom orchestrators, research-grade training stacks, or frontier model training operations.
- Experience with FinOps for AI workloads.
Culture & Benefits
- Full-time direct W2 employment.
- Remote work within the United States.
- Career growth opportunities within an established organization.
- Applicants must be U.S. citizens, Green Card holders, EAD holders, or H-1B transfer candidates; new H-1B visa petitions cannot be sponsored.
- Equal employment opportunity across legally protected statuses.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
5 дней назад
Principal Software Engineer (AI)
160 200 - 425 000$
Benzinga
2 дня назад
AI Engineer (Go)
90 000 - 100 000$
4 дня назад
Product Solutions Engineer (AI)
125 000 - 140 000$
14 часов назад
Senior AI Engineer
150 000 - 250 000$
3 дня назад
Principal AI Engineer
180 000 - 200 000$