Назад
Company hidden
2 дня назад

Data Center Operations Engineer (AI/HPC)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
Singapore/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Center Operations Engineer (AI/HPC): Operating and maintaining AI/HPC data center infrastructure, including NVIDIA GPU clusters, servers, storage, networking, and cabling systems with an accent on hardware maintenance, Linux administration, and high-speed interconnects. Focus on troubleshooting cluster failures, performing installations and upgrades, supporting infrastructure deployments, and responding to incidents in a 24x7 shift rotation.

Location: Needham, Massachusetts, United States

Company

hirify.global develops Bitcoin mining solutions and AI computational infrastructure, operating data centers across multiple countries.

What you will do

  • Operate and maintain data center infrastructure to ensure high availability and stable service operation.
  • Install, rack, cable, commission, maintain, and troubleshoot NVIDIA GPU clusters, AI/HPC servers, storage systems, and Ethernet and InfiniBand networks.
  • Monitor cluster health and perform hardware diagnostics, FRU replacement, BIOS/BMC/firmware upgrades, and preventive maintenance.
  • Provision servers, install operating systems, expand clusters, validate networks, and conduct burn-in testing.
  • Respond to incidents, document maintenance activities, prepare shift handover reports, and maintain SOPs and incident records.
  • Collaborate with engineering, network, and infrastructure teams on deployments and operational improvements.

Requirements

  • Bachelor’s degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or a related discipline.
  • Basic understanding of data center infrastructure, server hardware, CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
  • Knowledge of TCP/IP, Ethernet, VLAN, LACP, InfiniBand or RoCE, and structured cabling including DAC, AOC, optical fiber, MPO, and LC connectors.
  • Basic Linux administration skills, including system monitoring, systemctl, journalctl, dmesg, ip, ethtool, and shell scripting.
  • Willingness to work a 24x7 three-shift rotation, including nights, weekends, and holidays.
  • Strong teamwork, communication, ownership, problem-solving, and adherence to operational and safety procedures.

Nice to have

  • Experience in data center operations, hardware maintenance, AI/HPC infrastructure, or GPU clusters.
  • Familiarity with NVIDIA GB200 and GB300 systems and large-scale cluster environments.
  • Experience with high-speed networking, Slurm, Kubernetes, Prometheus, or Grafana.

Culture & Benefits

  • Work in a next-generation AI data center environment supporting large-scale NVIDIA clusters.
  • Participate in operational incident response and continuous infrastructure improvement.
  • Full-time position with a rotating 24x7 operations schedule.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →