2 дня назад
Data Center Site Manager / Supervisor (AI/HPC Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Center Site Manager / Supervisor (AI/HPC Infrastructure): Leading daily operations and 24x7 support for a mission-critical data center with an accent on AI/HPC clusters, GPU servers, networking, cabling, and operational governance. Focus on supervising Operations Engineers, coordinating incident response, maintaining infrastructure reliability, and driving hardware deployment and lifecycle management.
Location: Needham, Massachusetts, United States
Company
develops Bitcoin mining infrastructure, AI computational infrastructure, data centers, and cloud capabilities for high-demand artificial intelligence workloads.
What you will do
- Lead daily data center operations and maintain infrastructure availability, reliability, and operational excellence.
- Supervise Operations Engineers, including workforce planning, shift scheduling, task assignment, coaching, and performance management.
- Ensure 24x7 coverage and act as the primary escalation point for incidents, coordinating root cause analysis and resolution.
- Oversee AI/HPC infrastructure, including NVIDIA B300 clusters, GPU and x86 servers, storage, Ethernet and InfiniBand switches, and structured cabling.
- Coordinate hardware installation, rack and stack activities, commissioning, infrastructure expansion, and lifecycle management.
- Maintain SOPs, EOPs, preventive maintenance programs, operational KPIs, compliance, and continuous improvement initiatives.
Requirements
- Bachelor's degree or higher in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or a related discipline.
- At least 5 years of experience in data center operations, IT infrastructure, or HPC/AI infrastructure management.
- At least 2 years of team leadership or people management experience.
- Strong knowledge of Linux administration, server hardware troubleshooting, firmware and lifecycle management, Ethernet, InfiniBand, and structured cabling.
- Familiarity with NVIDIA GPU architecture, NVLink, NVSwitch, AI cluster deployment, monitoring tools, scripting, and automation.
- Willingness to provide hands-on operational support and participate in on-call duties, including shift coverage during shortages or critical incidents.
Culture & Benefits
- Full-time position in a mission-critical data center environment.
- 24x7 shift operations with on-call participation.
- Focus on operational excellence, accountability, teamwork, safety, security, and continuous improvement.
- Collaboration with engineering, network, facilities, and vendor teams.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Network Engineer (AI/HPC)
94 000 - 117 000$
7 дней назад
Network Engineer (AI Network & Security)
CoreWeave
6 дней назад
Data Center Manager (AI Infrastructure)
95 000 - 105 000$
CoreWeave
3 дня назад
Data Center Manager
95 000 - 105 000$
6 дней назад
Senior Data Center Technician (AI/HPC)
76 000 - 94 000$
6 дней назад
Data Center Technician (AI/HPC)
60 000 - 75 000$