Назад
Company hidden
обновлено 3 дня назад

GPU Infrastructure Engineer

150 000 - 300 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
GPU Infrastructure Engineer (Linux/Python/Go): Building reliable, production-ready infrastructure that turns bare-metal GPU servers into scalable compute for frontier AI workloads with an accent on provisioning, hardware lifecycle automation, and fleet health. Focus on designing recovery-safe workflows, integrating server readiness with Kubernetes and SLURM, and securing machines through credential handling, tenant isolation, and data sanitization.

Location: San Francisco or remote within the United States

Salary: $150,000–$300,000 per year plus equity incentives

Company

Builds an open superintelligence stack combining compute, environments, evaluations, secure sandboxes, training, and deployment for frontier AI teams.

What you will do

  • Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers.
  • Automate BIOS, BMC, NIC, GPU driver, and firmware configuration with staged rollouts and recovery paths.
  • Develop hardware inventory and lifecycle services covering machine identity, configuration, health, and readiness.
  • Create acceptance tests and burn-in workflows for GPUs, memory, storage, and interconnects.
  • Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems.
  • Build observability, quarantine, repair, re-provisioning, secure credential handling, tenant isolation, and data sanitization workflows.

Requirements

  • 3+ years of experience operating Linux servers or building bare-metal infrastructure automation in production.
  • Hands-on experience with PXE/iPXE, DHCP, image provisioning, and out-of-band management such as Redfish or IPMI.
  • Strong software engineering and debugging skills in Python, Go, or a comparable language, plus Bash.
  • Experience designing reliable automation that handles partial failures, retries, and configuration drift.
  • Knowledge of Linux boot, systemd, kernel and driver troubleshooting, OS image management, networking fundamentals, and configuration management.
  • Ability to own operational incidents and collaborate across hardware, networking, platform, and datacenter teams.

Nice to have

  • Experience with large NVIDIA GPU fleets, DGX/HGX platforms, or heterogeneous server vendors.
  • Experience with MAAS, Ironic, Tinkerbell, or similar provisioning systems.
  • Experience integrating Kubernetes or SLURM with node lifecycle management.
  • Experience with hardware qualification, automated burn-in, fleet health scoring, or open-source infrastructure tooling.

Culture & Benefits

  • Work directly with customers building foundation-model training and large-scale inference infrastructure.
  • Collaborate across engineering, hardware, networking, platform, and datacenter teams.
  • Contribute to open frontier AI infrastructure and systems operating at planetary scale.
  • Receive cash compensation plus equity incentives.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →