Назад
Company hidden
4 дня назад

Data Center Operations Engineer (AI/HPC)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
Singapore/US/Norway +4 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Center Operations Engineer (AI/HPC): Operating and maintaining AI/HPC data center infrastructure, including NVIDIA GB200 and GB300 clusters, GPU servers, storage, networking, and cabling with an accent on hardware reliability, cluster operations, and infrastructure troubleshooting. Focus on performing deployments, firmware and component maintenance, incident response, high-speed network validation, and 24x7 two-shift operations.

Location: Cyberjaya or Johor Bahru, Malaysia; onsite data center operations with a 24x7 two-shift rotation, including night shifts, weekends, and holidays.

Company

hirify.global develops Bitcoin mining and AI cloud solutions, including ASIC hardware, mining rigs, data centers, and high-performance computing infrastructure.

What you will do

  • Operate and maintain data center infrastructure to ensure high availability and stable service operation.
  • Install, rack, cable, commission, maintain, and troubleshoot NVIDIA GB200 and GB300 clusters, GPU servers, x86 servers, storage servers, and Ethernet and InfiniBand switches.
  • Monitor servers, GPUs, storage, networking devices, and related infrastructure, performing diagnostics, FRU replacement, and BIOS, BMC, and firmware upgrades.
  • Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
  • Respond to hardware and infrastructure incidents, complete preventive maintenance, and maintain logs, SOPs, handover reports, and incident documentation.
  • Collaborate with engineering, network, and infrastructure teams on deployments and operational improvements.

Requirements

  • Bachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or a related discipline.
  • Basic understanding of data center infrastructure, server hardware architecture, and components including CPU, memory, storage, GPU, BMC/IPMI, and firmware.
  • Knowledge of TCP/IP, Ethernet, VLAN, LACP, InfiniBand or RoCE, and structured cabling including DAC, AOC, optical fiber, MPO, and LC connectors.
  • Basic Linux administration skills, including system monitoring, systemctl, journalctl, dmesg, ip, ethtool, and shell scripting.
  • Willingness to work in a 24x7 two-shift rotation, including night shifts, weekends, and holidays.
  • Strong ownership, teamwork, communication, incident response, attention to detail, and adherence to operational and safety procedures.

Nice to have

  • Experience in data center operations, hardware maintenance, AI/HPC infrastructure, GPU clusters, or large-scale cluster environments.
  • Familiarity with NVIDIA GB200 and GB300 systems and high-speed networking technologies.
  • Experience with Slurm, Kubernetes, Prometheus, or Grafana.

Culture & Benefits

  • Inclusive environment that values authenticity and diverse backgrounds.
  • Opportunity to contribute to digital asset and next-generation AI data center infrastructure.
  • Personal accountability, autonomy, learning opportunities, training, and mentoring.
  • Welfare benefits and opportunities to work on new projects and systems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →