Назад
4 часа назад

Senior Network Production Engineer, AI Supercomputing

119 800 - 234 700$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Network Production Engineer, AI Supercomputing (Ethernet/RDMA): Operating and improving a multi-rail Ethernet backend network for frontier-scale AI supercomputers with an accent on network health, large-job reliability, and automated fleet operations. Focus on diagnosing failures across GPUs, NICs, switches, and distributed workloads, automating remediation, and reducing the impact of network faults on large training jobs.

Location: Mountain View, California, United States

Base pay: USD $119,800–$234,700 per year for Software Engineering IC4 or USD $142,800–$274,800 per year for Software Engineering IC5. Separate San Francisco Bay Area and New York City ranges also apply.

Company

Build and operate frontier-scale AI supercomputers used to train advanced machine-learning models.

What you will do

  • Own availability, performance, operational readiness, and service-level indicators for the multi-rail Ethernet backend network.
  • Respond to high-severity incidents and diagnose failures across GPUs, NICs, switches, cables, firmware, drivers, network operating systems, and distributed workloads.
  • Develop automated mitigation for path avoidance, node quarantine, workload relocation, and component remediation.
  • Investigate collective-performance degradation, stalls, timeouts, restarts, and model FLOPs utilization loss with training teams.
  • Build monitoring, validation, diagnosis, remediation, inventory, topology, and fleet-wide change-management automation.
  • Lead cross-functional root-cause investigations and create runbooks for reliable round-the-clock operations.

Requirements

  • Bachelor’s degree in Computer Science or a related technical discipline and 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, or equivalent experience.
  • Experience operating large-scale datacenter or high-performance computing networks in production.
  • Experience with Ethernet, RDMA, RoCEv2, congestion control, routing, load balancing, and lossless or near-lossless network design.
  • Experience diagnosing infrastructure failures across switches, NICs, hosts, drivers, firmware, and distributed applications.
  • Linux experience and programming in Python, Go, C++, or another systems-oriented language.
  • Experience building production monitoring or automation systems, analyzing telemetry, and leading high-severity incident response.

Nice to have

  • Experience with GPU training clusters, MRC or other multi-rail and multipath Ethernet transports, NCCL, CUDA, GPUDirect RDMA, or distributed training frameworks.
  • Familiarity with ConnectX-class NICs, Ethernet switch ASICs, SONiC, SAI, InfiniBand, Kubernetes, Slurm, and workload placement.
  • Experience with streaming telemetry, gNMI, Prometheus, Datadog, Kusto, eBPF, or equivalent observability systems.
  • Experience managing fleet-wide firmware, driver, or network operating-system changes.

Culture & Benefits

  • Participate in a global on-call rotation providing round-the-clock coverage.
  • Collaborate with infrastructure teams, cloud and datacenter operators, Azure Networking, NVIDIA, and hardware and network-software partners.
  • Benefits and additional compensation may be available depending on the role and location.

Hiring process

  • Applications are accepted on an ongoing basis until the position is filled.
  • The position will remain open for a minimum of five days.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →