Назад
Company hidden
4 дня назад

AI Infrastructure Engineer

100 000 - 160 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Infrastructure Engineer (GPU/ML Infrastructure): Designing and operating the platform layer for large-scale AI training and inference with an accent on GPU clusters, distributed training, scheduling, storage performance, and reliability. Focus on building accelerator resource-sharing systems, integrating ML frameworks, optimizing high-performance networking and storage, and implementing fault-tolerant workloads at scale.

Location: 100% remote within the United States

Salary: $100,000–$160,000 annually

Company

hirify.global is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

What you will do

  • Design and operate GPU and accelerator infrastructure across on-premises clusters, cloud-managed services, and hybrid configurations.
  • Build scheduling, queueing, resource-sharing, and developer tooling systems for large-scale ML workloads.
  • Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified AI platform.
  • Operate high-performance storage, data pipelines, RDMA, InfiniBand, NCCL, and high-bandwidth networking architectures.
  • Implement observability, checkpointing, restart, fault tolerance, security controls, isolation, and access management for multi-tenant infrastructure.
  • Drive automation, capacity planning, operational documentation, and cost optimization across compute, storage, and networking.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 10+ years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, collective communication, Linux internals, networking, and high-performance storage.
  • Must be authorized to work in the United States; eligible profiles include U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates. New H-1B visa petitions cannot be sponsored.

Nice to have

  • Experience operating InfiniBand or RDMA networking at scale.
  • Contributions to open-source ML infrastructure projects.
  • Experience with custom orchestrators, research-grade training stacks, frontier model training operations, or FinOps for AI workloads.

Culture & Benefits

  • Full-time direct W-2 employment.
  • Remote work within the United States.
  • Career growth opportunities within an established organization.
  • Equal employment opportunity and a workplace free from harassment and discrimination.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →