6 дней назад
GPU Compute & Bare Metal / DPU Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Compute & Bare Metal / DPU Engineer (GPU infrastructure): Operating the full lifecycle of bare-metal GPU nodes across a multi-region, 10,000+ GPU footprint with an accent on automated provisioning, firmware management, and fleet reliability. Focus on building node-delivery pipelines, managing DPU/SmartNIC and server firmware, reducing MTTR, and improving incident response and hardware-health monitoring.
Location: Remote within San Jose, CA or Austin, TX
Company
provides Bitcoin mining infrastructure and AI computational infrastructure, including data center design, equipment operations, and cloud capabilities for artificial intelligence workloads.
What you will do
- Own the full lifecycle of bare-metal GPU nodes, from provisioning and onboarding through operations, break-fix, and decommissioning across multiple regions.
- Build automated and repeatable node-delivery pipelines to support fleet growth toward 10,000+ GPUs.
- Manage DPU/SmartNIC and server firmware, including BMC, BIOS, NIC, and GPU version baselines, upgrades, and validation.
- Improve fleet reliability by reducing MTTR, leading incident response and root-cause analysis, and strengthening hardware-health monitoring.
- Participate in a sustainable 7×24 multi-region on-call rotation and create runbooks and tooling that reduce manual work.
- Partner with Storage/Image and Network teams to improve provisioning handoff and define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.
Requirements
- 3+ years of experience in large-scale bare-metal or server-fleet operations, HPC, or cloud infrastructure; 6+ years for the senior level.
- Hands-on experience operating GPU servers at scale, including NVIDIA HGX/DGX-class systems, driver, CUDA, and firmware management.
- Strong Linux systems skills with PXE, IPMI, Redfish, OS imaging, and automated provisioning.
- Experience with DPU/SmartNIC technologies such as NVIDIA BlueField and bare-metal networking.
- Infrastructure automation skills with Ansible, Terraform, Python, or Go.
- Experience with on-call operations, incident management, operational runbooks, and preferably multi-region or large-fleet operations.
Culture & Benefits
- Remote work is available within the specified US locations.
- Participation in a 7×24 multi-region on-call rotation is required.
- Work on AI and Bitcoin mining infrastructure spanning multiple data-center regions.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →