Назад
Company hidden
3 дня назад

Senior AI Cloud Network Operations Engineer (AI)

Тип работы
fulltime
Грейд
senior
Английский
c1
Страна
Singapore/Malaysia/Taiwan
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior AI Cloud Network Operations Engineer (AI): Monitoring and optimizing high-performance network infrastructure for AI cloud and Bitcoin mining with an accent on GPU-to-GPU communication and lossless network congestion control. Focus on resolving microsecond-level jitter, optimizing NCCL throughput, and managing large-scale GPU clusters.

Location: Singapore, Cyberjaya (MY), or Taibei (TW)

Company

World-leading technology company specializing in Bitcoin mining solutions and AI cloud infrastructure.

What you will do

  • Monitor 24/7 network estate including switches, routers, and optical transport to track health and traffic load.
  • Lead incident response for network outages and link failures to consistently hit SLA targets.
  • Troubleshoot high-performance AI network issues, specifically microsecond-level jitter and NCCL throughput degradation.
  • Manage the full ticket lifecycle for network requests, including fault reports and resource allocation.
  • Execute network changes under strict change management to minimize impact on AI training jobs.
  • Serve as the technical interface for internal R&D and customers for VLAN and routing policy adjustments.

Requirements

  • 5+ years of network operations experience in large-scale cloud, carrier NOC, or datacenter environments.
  • Deep knowledge of BGP, OSPF, VXLAN, EVPN, and ECMP.
  • Hands-on experience with NVIDIA Spectrum/Quantum, Arista, or Cisco CLI.
  • Working knowledge of InfiniBand, RoCEv2, and congestion control (PFC, ECN).
  • Proficiency with monitoring tools such as Zabbix, Prometheus, and Grafana.
  • Fluent in Chinese and English.

Nice to have

  • Experience operating GPU clusters with 1,000+ GPUs (NVIDIA H100 or GB200).
  • Familiarity with AI distributed training libraries such as NCCL or MPI.
  • Proficiency with NVIDIA UFM, including SHARP and network telemetry.
  • Professional certifications like CCIE, JNCIE, or NCP-AIN.
  • Mastery of Python or Go for developing network automation scripts.

Culture & Benefits

  • Inclusive environment that values authenticity and diversity of thought.
  • Fast-growing company offering a startup spirit and the chance to work with industrial pioneers.
  • High degree of personal accountability, autonomy, and opportunities for rapid professional growth.
  • Direct impact on the future of the digital asset and AI infrastructure industry.
  • Attractive welfare benefits, including dedicated mentoring and training programs.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →