Назад
Company hidden
25 дней назад

Sr. GPU Cloud East-West Network Expert (SRE SME)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. GPU Cloud East-West Network Expert (SRE SME) (InfiniBand/RoCE): Operating and optimizing high-scale GPU cluster networks across InfiniBand and RoCEv2 fabrics with an accent on topology design, RDMA performance, and fabric reliability. Focus on feeding telemetry into AIOps models, diagnosing complex network faults, and converting incident mitigations into automated remediation workflows.

Location: Singapore, SG / Penang, Malaysia

Company

hirify.global provides Bitcoin mining solutions, AI cloud capabilities, ASIC hardware, and datacenter infrastructure and operations.

What you will do

  • Design, deploy, operate, and optimize InfiniBand fabrics using fat-tree, rail-optimized, and dragonfly topologies for GPU clusters ranging from 100 to 10,000 GPUs.
  • Manage RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • Monitor and tune IB and RoCE performance, including adaptive routing, DCQCN/ECN congestion control, traffic isolation, and NCCL communication.
  • Manage UFM-based fabric monitoring, diagnostics, subnet management, firmware lifecycles, and fault investigation.
  • Diagnose link flaps, symbol errors, packet drops, routing anomalies, and credit stalls, coordinating escalations and RMAs with Nvidia/Mellanox.
  • Feed IB/RoCE telemetry into the AIOps platform and convert incidents and mitigations into labeled training data and automated remediation workflows.

Requirements

  • 5+ years of data center networking experience, including at least 3 years focused on InfiniBand or RoCE fabrics.
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale.
  • Strong knowledge of IB subnet management, partitioning, QoS, and RoCEv2 configuration with PFC, ECN, and DCQCN.
  • Proficiency with UFM or equivalent InfiniBand fabric management tools and diagnostic utilities such as ibdiagnet, perfquery, and ibstat.
  • Knowledge of 400G/800G optics, cabling standards, structured cabling, NCCL, and GPU-to-network topology mapping.
  • Experience with telemetry-driven operations and a runbook-as-code approach to automation.

Culture & Benefits

  • Inclusive environment that values authenticity and diverse perspectives.
  • Startup-oriented setting in a fast-growing company.
  • Opportunity to contribute to digital asset and AI cloud infrastructure projects.
  • Autonomy, personal accountability, and opportunities for growth and learning.
  • Training, mentoring, and welfare benefits.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →