Назад
Company hidden
обновлено 20 часов назад

Datacenter Operations Engineer (AI Infrastructure)

150 000 - 300 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Datacenter Operations Engineer (AI Infrastructure) (GPU cloud): Operating and scaling the physical infrastructure behind a GPU cloud, with an accent on rack deployments, hardware maintenance, incident response, and high-density datacenter systems. Focus on coordinating remote hands and vendors, troubleshooting server and networking hardware, automating operational workflows, and reducing repair times and customer disruption.

Location: San Francisco; Remote

Salary: $150,000–$300,000 per year, plus equity incentives

Company

hirify.global is building an open superintelligence stack that provides AI teams with compute, environments, evaluations, secure sandboxes, training, and deployment infrastructure.

What you will do

  • Coordinate rack deployments, cabling, inventory, acceptance testing, and new GPU capacity with datacenter partners and engineering teams.
  • Maintain asset records, rack layouts, power allocations, cabling documentation, and spare-parts inventories.
  • Lead hardware fault triage, remote-hands coordination, vendor escalations, component replacement, and RMA workflows.
  • Establish maintenance plans, change procedures, runbooks, and escalation processes that minimize customer disruption.
  • Track capacity readiness, hardware failures, repair times, and operational risks while automating reporting and workflows.
  • Partner with facility teams on power, cooling, environmental monitoring, and high-density GPU deployment readiness.

Requirements

  • 3+ years of experience in datacenter operations, hardware infrastructure, or production systems operations.
  • Hands-on experience deploying and troubleshooting rack-mounted servers, networking equipment, and structured cabling.
  • Experience coordinating datacenter providers, remote hands, and hardware vendors during deployments and incidents.
  • Working knowledge of Linux diagnostics, BMC consoles, and server hardware health tools.
  • Knowledge of GPU server components, PCIe devices, memory, storage, hardware diagnostics, rack power, redundant power paths, airflow, and high-density cooling.
  • Experience with fiber and copper cabling, optics, labeling, asset tracking, spares management, change control, incident management, and basic scripting.

Nice to have

  • Experience with NVIDIA DGX/HGX systems or large GPU cluster deployments.
  • Experience with liquid-cooled infrastructure and facility engineering teams.
  • Experience bringing up new datacenter sites or expanding multi-site capacity.
  • Experience with hardware qualification, burn-in testing, reliability analysis, or automated fleet provisioning.

Culture & Benefits

  • Work directly with customers training foundation models and deploying large-scale inference infrastructure.
  • Collaborate with an engineering team building infrastructure for frontier AI research and production systems.
  • Have direct impact on reliable, high-performance GPU infrastructure and large-scale AI deployments.
  • Receive cash compensation plus equity incentives.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →