9 дней назад
Data Center Site Manager / Supervisor (AI/HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Center Site Manager / Supervisor (AI/HPC): Leading daily operations and infrastructure management at an AI and HPC data center site with an accent on 24x7 availability, NVIDIA GPU clusters, Linux administration, and operational governance. Focus on supervising Operations Engineers, coordinating incident resolution, expanding infrastructure, and maintaining reliability across mission-critical systems.
Location: Aurora, Colorado, United States
Company
develops Bitcoin mining solutions and AI computational infrastructure, managing data center design, construction, equipment, cloud capabilities, and daily operations across multiple countries.
What you will do
- Lead daily operations of the data center site and maintain infrastructure availability, reliability, and operational excellence.
- Supervise Operations Engineers through workforce planning, shift scheduling, task assignment, coaching, and performance management.
- Ensure 24x7 operational coverage and serve as the primary escalation point for incidents, coordinating root cause analysis and resolution.
- Oversee AI/HPC infrastructure, including NVIDIA B300 clusters, GPU servers, x86 servers, storage, Ethernet and InfiniBand networking, and structured cabling.
- Coordinate hardware installation, rack and stack activities, system commissioning, infrastructure expansion, and lifecycle management.
- Develop and improve SOPs, EOPs, preventive maintenance programs, operational KPIs, safety procedures, and security controls.
Requirements
- Bachelor’s degree or higher in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or a related discipline.
- At least 5 years of experience in data center operations, IT infrastructure, or HPC/AI infrastructure management.
- At least 2 years of experience leading teams or managing people, including shift scheduling, incident escalation, coaching, and vendor coordination.
- Strong knowledge of Linux administration, hardware troubleshooting, firmware management, performance diagnostics, log analysis, networking, and basic scripting or automation.
- Experience with GPU and x86 servers, storage systems, Ethernet and InfiniBand networking, NVIDIA GPU architecture, NVLink, NVSwitch, optical fiber, DAC, and AOC cabling.
- Willingness to participate in on-call duties and provide hands-on support during shortages, emergencies, critical incidents, or shift coverage.
Nice to have
- Experience managing 24x7 shift operations in a mission-critical environment.
- Experience with large-scale AI or HPC clusters.
- Familiarity with infrastructure monitoring and management tools.
Culture & Benefits
- Focus on operational excellence, teamwork, accountability, and continuous improvement.
- Hands-on work in AI computational infrastructure and Bitcoin mining technology.
- Collaboration with engineering, network, facilities, and vendor teams.
- Equal employment opportunities are provided in accordance with applicable country, state, and local laws.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Network Engineer (AI/HPC)
94 000 - 117 000$
13 дней назад
Data Center Technician (AI/HPC)
60 000 - 75 000$
13 дней назад
Senior Data Center Technician (AI/HPC)
76 000 - 94 000$
CoreWeave
13 дней назад
Data Center Manager (AI Infrastructure)
95 000 - 105 000$
CoreWeave
10 дней назад
Data Center Manager
95 000 - 105 000$
13 дней назад