Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Head of Infrastructure Support (AI): Leading the regional infrastructure support operations for high-performance GPU superclusters with an accent on service reliability, team scaling, and complex incident management. Focus on building a high-performing APAC support team, driving operational excellence, and ensuring seamless follow-the-sun service delivery for AI-native enterprises.
Location: Must be based in Singapore
Company
Nscale is a vertically integrated AI cloud provider delivering high-performance infrastructure, including GPU superclusters and orchestration, to AI-native companies and enterprises.
What you will do
- Own regional service outcomes, including SLA adherence, MTTR, and customer satisfaction for the APAC region.
- Manage the full lifecycle of the regional support team, including hiring, performance reviews, and professional development.
- Partner with global counterparts to establish consistent standards, processes, and follow-the-sun coverage models.
- Act as the senior escalation point for complex, high-impact incidents and lead post-incident reviews.
- Drive capacity modelling and headcount planning to support rapid regional growth.
- Maintain technical credibility by guiding complex troubleshooting on GPU infrastructure, Linux systems, and high-performance fabrics.
Requirements
- 5+ years of direct line management experience in an operational support environment with end-to-end performance management.
- 8+ years of technical experience in Linux systems engineering, GPU infrastructure, and high-performance networking.
- Proven expertise in HPC scheduling (Slurm), RDMA fabrics (InfiniBand/RoCE), and data centre operations.
- Strong understanding of ITIL-aligned incident, problem, and change management practices.
- Excellent written and verbal communication skills, capable of presenting to both engineers and executives.
- Must be based in Singapore with the ability to participate in regional on-call and travel as required.
Nice to have
- Deeper exposure to NCCL, NVLink, or AI-optimized storage platforms like VAST or Ceph.
- Experience with OpenStack or fleet-scale provisioning tools like MAAS or NetBox.
- Previous experience running multi-region or follow-the-sun support operations.
- Relevant certifications in ITIL, Linux, or cloud infrastructure.
Culture & Benefits
- Culture of relentless innovation, ownership, and accountability.
- Commitment to transparency, openness, and customer-centric focus.
- Fast-paced, collaborative environment with high standards for excellence.
- Focus on sustainability and responsible technology development.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Data Center Technician (AI)
80 000 - 90 000$
5 дней назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
3 дня назад
Infrastructure Engineer (Storage)
180 000 - 220 000$
10 часов назад
Head of Network Engineering (AI)
176 000 - 221 000$
5 дней назад
Systems Engineer (AWS/Linux)
5 дней назад