12 дней назад
Senior Platform Reliability Engineer (Fabric and Interconnect)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Platform Reliability Engineer (Fabric and Interconnect) (AI infrastructure): Operating and improving GPU-to-GPU interconnect domains, high-performance network fabrics, and DPU-based host networking for large-scale AI infrastructure with an accent on automation, fault diagnosis, and production reliability. Focus on building self-healing remediation, tuning collective communication and congestion control, executing firmware and capacity changes, and leading complex fabric incident recovery.
Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required
Company
develops and operates energy-efficient AI infrastructure, including AI Cloud and the proprietary AI FactoryOS platform.
What you will do
- Own the reliable operation, automation, and continuous improvement of GPU interconnect and high-performance network fabrics, including NVLink, NVSwitch, InfiniBand, and Spectrum-X.
- Build guarded automation and remediation tooling for fault isolation, fabric reconvergence, and software-driven recovery of AI clusters.
- Diagnose and tune performance across the interconnect stack, from application collective communication to physical link level, including RoCE and congestion control.
- Operate DPU-based host networking, including offload paths and driver and firmware compatibility.
- Plan and execute firmware upgrade waves, fabric expansions, capacity changes, staged rollouts, production verification, and rollback procedures.
- Lead major fabric incident recovery and vendor escalations, drive permanent fixes, and mentor engineers through runbooks and technical guidance.
Requirements
- 8+ years of experience in high-performance networking and systems engineering, including ownership of production network or interconnect infrastructure in a 24/7 environment.
- Deep operational experience with GPU interconnect fabrics, including NVLink and NVSwitch topology and failure diagnosis.
- Extensive experience with InfiniBand or RoCE-based Ethernet fabrics, routing internals, and congestion control tuning.
- Experience with DPU or SmartNIC-based host networking, offload paths, and driver and firmware coordination.
- Strong infrastructure automation and infrastructure-as-code skills, plus practical scripting or programming experience with Python, Go, or Bash.
- Experience with major incident response, on-call participation, engineering-level vendor escalation, production change control, and executable runbooks.
Nice to have
- Experience operating fabrics for large-scale distributed training or inference workloads.
- Experience with NVIDIA rack-scale or multi-node GPU systems and interconnect topology.
- Experience in a multi-tenant service provider, cloud, or colocation environment.
- Knowledge of data centre and hardware fundamentals, including cabling, optics, firmware management, and hardware fault workflows.
- Relevant vendor certification.
Culture & Benefits
- Permanent full-time employment.
- Participation in a shared after-hours escalation roster supporting a 24/7 function.
- High degree of autonomy, broad technical direction, and direct access to decision makers.
- Work focused on sustainable AI infrastructure, efficient computing, and large-scale GPU operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Senior Site Reliability Engineer (AI)
14 дней назад
Platform Support Engineer (Trading)
13 дней назад
Senior/Lead Site Reliability Engineer (AI/LLM)
Airwallex
13 дней назад
Senior Site Reliability Engineer (Fintech)
12 дней назад
Senior Data Site Reliability Engineer (Video Platform)
12 дней назад