Назад
Company hidden
2 дня назад

Data Center Site Manager / Supervisor (AI/HPC Infrastructure)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Data Center Site Manager / Supervisor (AI/HPC Infrastructure): Leading daily operations and 24x7 support for a mission-critical data center with an accent on AI/HPC clusters, GPU servers, networking, cabling, and operational governance. Focus on supervising Operations Engineers, coordinating incident response, maintaining infrastructure reliability, and driving hardware deployment and lifecycle management.

Location: Needham, Massachusetts, United States

Company

hirify.global develops Bitcoin mining infrastructure, AI computational infrastructure, data centers, and cloud capabilities for high-demand artificial intelligence workloads.

What you will do

  • Lead daily data center operations and maintain infrastructure availability, reliability, and operational excellence.
  • Supervise Operations Engineers, including workforce planning, shift scheduling, task assignment, coaching, and performance management.
  • Ensure 24x7 coverage and act as the primary escalation point for incidents, coordinating root cause analysis and resolution.
  • Oversee AI/HPC infrastructure, including NVIDIA B300 clusters, GPU and x86 servers, storage, Ethernet and InfiniBand switches, and structured cabling.
  • Coordinate hardware installation, rack and stack activities, commissioning, infrastructure expansion, and lifecycle management.
  • Maintain SOPs, EOPs, preventive maintenance programs, operational KPIs, compliance, and continuous improvement initiatives.

Requirements

  • Bachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or a related discipline.
  • At least 5 years of experience in data center operations, IT infrastructure, or HPC/AI infrastructure management.
  • At least 2 years of team leadership or people management experience.
  • Strong knowledge of Linux administration, server hardware troubleshooting, firmware and lifecycle management, Ethernet, InfiniBand, and structured cabling.
  • Familiarity with NVIDIA GPU architecture, NVLink, NVSwitch, AI cluster deployment, monitoring tools, scripting, and automation.
  • Willingness to provide hands-on operational support and participate in on-call duties, including shift coverage during shortages or critical incidents.

Culture & Benefits

  • Full-time position in a mission-critical data center environment.
  • 24x7 shift operations with on-call participation.
  • Focus on operational excellence, accountability, teamwork, safety, security, and continuous improvement.
  • Collaboration with engineering, network, facilities, and vendor teams.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →