2 дня назад
Data Center Operations Engineer (AI/HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Data Center Operations Engineer (AI/HPC): Operating and maintaining AI/HPC data center infrastructure, including NVIDIA GPU clusters, servers, storage, networking, and cabling systems with an accent on hardware maintenance, Linux administration, and high-speed interconnects. Focus on troubleshooting cluster failures, performing installations and upgrades, supporting infrastructure deployments, and responding to incidents in a 24x7 shift rotation.
Location: Needham, Massachusetts, United States
Company
develops Bitcoin mining solutions and AI computational infrastructure, operating data centers across multiple countries.
What you will do
- Operate and maintain data center infrastructure to ensure high availability and stable service operation.
- Install, rack, cable, commission, maintain, and troubleshoot NVIDIA GPU clusters, AI/HPC servers, storage systems, and Ethernet and InfiniBand networks.
- Monitor cluster health and perform hardware diagnostics, FRU replacement, BIOS/BMC/firmware upgrades, and preventive maintenance.
- Provision servers, install operating systems, expand clusters, validate networks, and conduct burn-in testing.
- Respond to incidents, document maintenance activities, prepare shift handover reports, and maintain SOPs and incident records.
- Collaborate with engineering, network, and infrastructure teams on deployments and operational improvements.
Requirements
- Bachelor’s degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or a related discipline.
- Basic understanding of data center infrastructure, server hardware, CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
- Knowledge of TCP/IP, Ethernet, VLAN, LACP, InfiniBand or RoCE, and structured cabling including DAC, AOC, optical fiber, MPO, and LC connectors.
- Basic Linux administration skills, including system monitoring, systemctl, journalctl, dmesg, ip, ethtool, and shell scripting.
- Willingness to work a 24x7 three-shift rotation, including nights, weekends, and holidays.
- Strong teamwork, communication, ownership, problem-solving, and adherence to operational and safety procedures.
Nice to have
- Experience in data center operations, hardware maintenance, AI/HPC infrastructure, or GPU clusters.
- Familiarity with NVIDIA GB200 and GB300 systems and large-scale cluster environments.
- Experience with high-speed networking, Slurm, Kubernetes, Prometheus, or Grafana.
Culture & Benefits
- Work in a next-generation AI data center environment supporting large-scale NVIDIA clusters.
- Participate in operational incident response and continuous infrastructure improvement.
- Full-time position with a rotating 24x7 operations schedule.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Data Center Technician (AI/HPC)
60 000 - 75 000$
6 дней назад
Senior Data Center Technician (AI/HPC)
76 000 - 94 000$
CoreWeave
8 дней назад
Data Center Technician (AI Infrastructure)
65 000 - 83 000$
CoreWeave
6 дней назад
Data Center Manager (AI Infrastructure)
95 000 - 105 000$
Lambda
3 дня назад
Data Center Technician (AI)
89 000 - 119 000$
6 дней назад
Storage Engineer (AI/HPC)
94 000 - 117 000$