обновлено 3 дня назад
GPU Infrastructure Engineer
150 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
GPU Infrastructure Engineer (Linux/Python/Go): Building reliable, production-ready infrastructure that turns bare-metal GPU servers into scalable compute for frontier AI workloads with an accent on provisioning, hardware lifecycle automation, and fleet health. Focus on designing recovery-safe workflows, integrating server readiness with Kubernetes and SLURM, and securing machines through credential handling, tenant isolation, and data sanitization.
Location: San Francisco or remote within the United States
Salary: $150,000–$300,000 per year plus equity incentives
Company
Builds an open superintelligence stack combining compute, environments, evaluations, secure sandboxes, training, and deployment for frontier AI teams.
What you will do
- Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers.
- Automate BIOS, BMC, NIC, GPU driver, and firmware configuration with staged rollouts and recovery paths.
- Develop hardware inventory and lifecycle services covering machine identity, configuration, health, and readiness.
- Create acceptance tests and burn-in workflows for GPUs, memory, storage, and interconnects.
- Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems.
- Build observability, quarantine, repair, re-provisioning, secure credential handling, tenant isolation, and data sanitization workflows.
Requirements
- 3+ years of experience operating Linux servers or building bare-metal infrastructure automation in production.
- Hands-on experience with PXE/iPXE, DHCP, image provisioning, and out-of-band management such as Redfish or IPMI.
- Strong software engineering and debugging skills in Python, Go, or a comparable language, plus Bash.
- Experience designing reliable automation that handles partial failures, retries, and configuration drift.
- Knowledge of Linux boot, systemd, kernel and driver troubleshooting, OS image management, networking fundamentals, and configuration management.
- Ability to own operational incidents and collaborate across hardware, networking, platform, and datacenter teams.
Nice to have
- Experience with large NVIDIA GPU fleets, DGX/HGX platforms, or heterogeneous server vendors.
- Experience with MAAS, Ironic, Tinkerbell, or similar provisioning systems.
- Experience integrating Kubernetes or SLURM with node lifecycle management.
- Experience with hardware qualification, automated burn-in, fleet health scoring, or open-source infrastructure tooling.
Culture & Benefits
- Work directly with customers building foundation-model training and large-scale inference infrastructure.
- Collaborate across engineering, hardware, networking, platform, and datacenter teams.
- Contribute to open frontier AI infrastructure and systems operating at planetary scale.
- Receive cash compensation plus equity incentives.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Motive
10 дней назад
Staff Engineer (Developer Productivity)
164 000 - 236 000$
11 дней назад
Staff Platform Engineer (AI/ML)
126 000 - 174 000$
11 дней назад
Staff Platform Engineer (AI/ML)
103 000 - 143 000CAD
11 дней назад
Senior Platform Engineer (AI)
228 000 - 279 000$
11 дней назад
Principal ML Platform Engineer (AI)
175 000 - 325 000$
11 дней назад
Senior Software Engineer (Infrastructure)
75 000 - 100 000€