4 часа назад
Senior Network Production Engineer, AI Supercomputing
119 800 - 234 700$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Network Production Engineer, AI Supercomputing (Ethernet/RDMA): Operating and improving a multi-rail Ethernet backend network for frontier-scale AI supercomputers with an accent on network health, large-job reliability, and automated fleet operations. Focus on diagnosing failures across GPUs, NICs, switches, and distributed workloads, automating remediation, and reducing the impact of network faults on large training jobs.
Location: Mountain View, California, United States
Base pay: USD $119,800–$234,700 per year for Software Engineering IC4 or USD $142,800–$274,800 per year for Software Engineering IC5. Separate San Francisco Bay Area and New York City ranges also apply.
Company
Build and operate frontier-scale AI supercomputers used to train advanced machine-learning models.
What you will do
- Own availability, performance, operational readiness, and service-level indicators for the multi-rail Ethernet backend network.
- Respond to high-severity incidents and diagnose failures across GPUs, NICs, switches, cables, firmware, drivers, network operating systems, and distributed workloads.
- Develop automated mitigation for path avoidance, node quarantine, workload relocation, and component remediation.
- Investigate collective-performance degradation, stalls, timeouts, restarts, and model FLOPs utilization loss with training teams.
- Build monitoring, validation, diagnosis, remediation, inventory, topology, and fleet-wide change-management automation.
- Lead cross-functional root-cause investigations and create runbooks for reliable round-the-clock operations.
Requirements
- Bachelor’s degree in Computer Science or a related technical discipline and 4+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering, or equivalent experience.
- Experience operating large-scale datacenter or high-performance computing networks in production.
- Experience with Ethernet, RDMA, RoCEv2, congestion control, routing, load balancing, and lossless or near-lossless network design.
- Experience diagnosing infrastructure failures across switches, NICs, hosts, drivers, firmware, and distributed applications.
- Linux experience and programming in Python, Go, C++, or another systems-oriented language.
- Experience building production monitoring or automation systems, analyzing telemetry, and leading high-severity incident response.
Nice to have
- Experience with GPU training clusters, MRC or other multi-rail and multipath Ethernet transports, NCCL, CUDA, GPUDirect RDMA, or distributed training frameworks.
- Familiarity with ConnectX-class NICs, Ethernet switch ASICs, SONiC, SAI, InfiniBand, Kubernetes, Slurm, and workload placement.
- Experience with streaming telemetry, gNMI, Prometheus, Datadog, Kusto, eBPF, or equivalent observability systems.
- Experience managing fleet-wide firmware, driver, or network operating-system changes.
Culture & Benefits
- Participate in a global on-call rotation providing round-the-clock coverage.
- Collaborate with infrastructure teams, cloud and datacenter operators, Azure Networking, NVIDIA, and hardware and network-software partners.
- Benefits and additional compensation may be available depending on the role and location.
Hiring process
- Applications are accepted on an ongoing basis until the position is filled.
- The position will remain open for a minimum of five days.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Datadog
4 дня назад
Senior Software Engineer - Incident Insights & Readiness (SRE)
192 000 - 240 000$
4 дня назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
152 000 - 228 000$
6 дней назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
3 дня назад
Senior Incident Commander (SRE/Python)
187 000 - 233 500$
4 дня назад
Senior Staff Engineer (SRE)
130 000 - 260 000$
6 дней назад
Senior Staff Engineer (SRE)
120 000 - 260 000$