Назад
Company hidden
2 дня назад

Infrastructure, Large-scale Training (AI)

180 000 - 450 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Infrastructure, Large-scale Training (AI) (GPU Clusters and ML Infrastructure): Building and operating large-scale GPU computing clusters for AI training and inference with an accent on reliability, scalability, and cost efficiency. Focus on provisioning 10,000+ GPU environments, optimizing networking and scheduling, and solving complex distributed-systems and production-reliability challenges.

Location: San Jose, United States

Salary: $180,000–$450,000 annually

Company

hirify.global is an artificial intelligence company developing proactive, multimodal intelligence and next-generation hardware that interact with people and the real world through speech, text, vision, and persistent memory.

What you will do

  • Design and maintain Infrastructure as Code practices for repeatable, auditable, and scalable GPU cluster provisioning.
  • Improve and harden CI/CD pipelines for secure, reliable, low-latency model delivery in production.
  • Own training infrastructure operating at the scale of 10,000+ GPUs, including scheduling, fault tolerance, and network fabric optimization.
  • Partner with ML researchers and engineers to identify compute bottlenecks and deliver infrastructure improvements.
  • Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.
  • Drive capacity planning, cost efficiency, hardware lifecycle management, and internal tooling for compute users.

Requirements

  • 5+ years of experience in infrastructure, systems, or platform engineering, including at least 2 years in ML or HPC environments.
  • Experience managing GPU clusters or large-scale distributed compute infrastructure.
  • Strong proficiency in at least one systems or infrastructure programming language.
  • Deep understanding of networking fundamentals relevant to high-throughput training workloads.
  • Experience with container orchestration, job scheduling, and multi-tenant resource management.
  • Production systems ownership, high-reliability operations, debugging, and observability across the infrastructure stack.

Nice to have

  • Experience operating large, GPU-aware Kubernetes clusters.
  • Pulumi or similar Infrastructure as Code tooling.
  • Rust or Go for systems-level tooling and performance-critical services.
  • Familiarity with PyTorch and Ray.
  • RDMA, InfiniBand, or RoCE experience.

Culture & Benefits

  • Full-time position with a US annual base salary range of $180,000–$450,000.
  • Total compensation may include additional components and benefits depending on the role.
  • Work centers on highly technical infrastructure-as-a-product engineering for AI research and production workloads.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →