Назад
Company hidden
4 дня назад

ML Infrastructure Engineer (AI)

100 000 - 150 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Infrastructure Engineer (AI): Designing and operating GPU infrastructure and platform systems for large-scale AI training and inference with an accent on distributed training, scheduling, storage performance, and reliability. Focus on integrating ML frameworks, building high-performance networking and observability, optimizing infrastructure costs, and implementing fault tolerance for multi-tenant workloads.

Location: 100% remote within the United States

Salary: $100,000–$150,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and operate GPU and accelerator infrastructure across on-premises clusters, cloud-managed services, and hybrid configurations.
  • Build scheduling, queueing, resource-sharing, and automation systems to improve accelerator utilization across teams.
  • Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, Ray Train, and related ML training frameworks.
  • Operate high-performance storage, data pipelines, RDMA/InfiniBand networking, NCCL, and collective communication systems.
  • Build observability, checkpointing, restart, fault-tolerance, security, and isolation capabilities for multi-tenant AI workloads.
  • Develop researcher tooling, capacity planning processes, operational documentation, and cost-optimization strategies.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 6+ years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong Python skills and proficiency in at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, collective communication, Linux internals, networking, and high-performance storage.
  • Experience with Kubernetes, Slurm, Ray, a major cloud provider’s ML infrastructure, testing, CI/CD, and code review.

Nice to have

  • Experience operating InfiniBand or RDMA networking at scale.
  • Open-source ML infrastructure contributions or familiarity with custom orchestrators and research-grade training stacks.
  • Exposure to frontier model training operations and FinOps for AI workloads.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Career growth opportunity within an established organization.
  • Remote work arrangement within the United States.
  • New H-1B visa petitions are not sponsored; U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates may apply.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →