25 дней назад
Sr. GPU Cloud East-West Network Expert (SRE SME)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. GPU Cloud East-West Network Expert (SRE SME) (InfiniBand/RoCE): Operating and optimizing high-scale GPU cluster networks across InfiniBand and RoCEv2 fabrics with an accent on topology design, RDMA performance, and fabric reliability. Focus on feeding telemetry into AIOps models, diagnosing complex network faults, and converting incident mitigations into automated remediation workflows.
Location: Singapore, SG / Penang, Malaysia
Company
provides Bitcoin mining solutions, AI cloud capabilities, ASIC hardware, and datacenter infrastructure and operations.
What you will do
- Design, deploy, operate, and optimize InfiniBand fabrics using fat-tree, rail-optimized, and dragonfly topologies for GPU clusters ranging from 100 to 10,000 GPUs.
- Manage RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
- Monitor and tune IB and RoCE performance, including adaptive routing, DCQCN/ECN congestion control, traffic isolation, and NCCL communication.
- Manage UFM-based fabric monitoring, diagnostics, subnet management, firmware lifecycles, and fault investigation.
- Diagnose link flaps, symbol errors, packet drops, routing anomalies, and credit stalls, coordinating escalations and RMAs with Nvidia/Mellanox.
- Feed IB/RoCE telemetry into the AIOps platform and convert incidents and mitigations into labeled training data and automated remediation workflows.
Requirements
- 5+ years of data center networking experience, including at least 3 years focused on InfiniBand or RoCE fabrics.
- Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale.
- Strong knowledge of IB subnet management, partitioning, QoS, and RoCEv2 configuration with PFC, ECN, and DCQCN.
- Proficiency with UFM or equivalent InfiniBand fabric management tools and diagnostic utilities such as ibdiagnet, perfquery, and ibstat.
- Knowledge of 400G/800G optics, cabling standards, structured cabling, NCCL, and GPU-to-network topology mapping.
- Experience with telemetry-driven operations and a runbook-as-code approach to automation.
Culture & Benefits
- Inclusive environment that values authenticity and diverse perspectives.
- Startup-oriented setting in a fast-growing company.
- Opportunity to contribute to digital asset and AI cloud infrastructure projects.
- Autonomy, personal accountability, and opportunities for growth and learning.
- Training, mentoring, and welfare benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 дней назад
Senior Platform Reliability Engineer (Fabric and Interconnect)
11 дней назад
Sr. Site Reliability Engineer (AI)
13 дней назад
Senior Database Reliability Engineer (AI)
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
9 дней назад
Head of Site Reliability Engineering (AI)
195 000 - 285 000$
12 дней назад
Senior Cloud Infrastructure and Networking
125 000 - 135 000$