Назад
Company hidden
12 дней назад

Senior Platform Reliability Engineer (Fabric and Interconnect)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Platform Reliability Engineer (Fabric and Interconnect) (AI infrastructure): Operating and improving GPU-to-GPU interconnect domains, high-performance network fabrics, and DPU-based host networking for large-scale AI infrastructure with an accent on automation, fault diagnosis, and production reliability. Focus on building self-healing remediation, tuning collective communication and congestion control, executing firmware and capacity changes, and leading complex fabric incident recovery.

Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required

Company

hirify.global develops and operates energy-efficient AI infrastructure, including hirify.global AI Cloud and the proprietary AI FactoryOS platform.

What you will do

  • Own the reliable operation, automation, and continuous improvement of GPU interconnect and high-performance network fabrics, including NVLink, NVSwitch, InfiniBand, and Spectrum-X.
  • Build guarded automation and remediation tooling for fault isolation, fabric reconvergence, and software-driven recovery of AI clusters.
  • Diagnose and tune performance across the interconnect stack, from application collective communication to physical link level, including RoCE and congestion control.
  • Operate DPU-based host networking, including offload paths and driver and firmware compatibility.
  • Plan and execute firmware upgrade waves, fabric expansions, capacity changes, staged rollouts, production verification, and rollback procedures.
  • Lead major fabric incident recovery and vendor escalations, drive permanent fixes, and mentor engineers through runbooks and technical guidance.

Requirements

  • 8+ years of experience in high-performance networking and systems engineering, including ownership of production network or interconnect infrastructure in a 24/7 environment.
  • Deep operational experience with GPU interconnect fabrics, including NVLink and NVSwitch topology and failure diagnosis.
  • Extensive experience with InfiniBand or RoCE-based Ethernet fabrics, routing internals, and congestion control tuning.
  • Experience with DPU or SmartNIC-based host networking, offload paths, and driver and firmware coordination.
  • Strong infrastructure automation and infrastructure-as-code skills, plus practical scripting or programming experience with Python, Go, or Bash.
  • Experience with major incident response, on-call participation, engineering-level vendor escalation, production change control, and executable runbooks.

Nice to have

  • Experience operating fabrics for large-scale distributed training or inference workloads.
  • Experience with NVIDIA rack-scale or multi-node GPU systems and interconnect topology.
  • Experience in a multi-tenant service provider, cloud, or colocation environment.
  • Knowledge of data centre and hardware fundamentals, including cabling, optics, firmware management, and hardware fault workflows.
  • Relevant vendor certification.

Culture & Benefits

  • Permanent full-time employment.
  • Participation in a shared after-hours escalation roster supporting a 24/7 function.
  • High degree of autonomy, broad technical direction, and direct access to decision makers.
  • Work focused on sustainable AI infrastructure, efficient computing, and large-scale GPU operations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →