Senior Infrastructure Engineer (GPU Compute)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Infrastructure Engineer (GPU Compute): Architecting, deploying, and operating production GPU and HPC clusters for a liquid GPU offtake marketplace with an accent on hardware performance, reliability, and fleet automation. Focus on designing large-scale compute infrastructure, troubleshooting issues across hardware and software layers, and building provisioning, monitoring, and remediation systems.
Location: San Francisco, CA; hybrid work with 3–4 days per week in the office and domestic travel when required. Visa and work permit sponsorship is available.
Salary: $220,000–$300,000 plus equity.
Company
is building a liquid market for GPU offtake and infrastructure supporting large-scale AI compute.
What you will do
- Architect and deploy new GPU and HPC clusters around the world.
- Operate production compute infrastructure and participate in the on-call rotation.
- Deploy environments, resolve hardware and software incidents, and improve system reliability.
- Automate provisioning, monitoring, remediation, and fleet operations at scale.
- Document operational procedures and maintain runbooks.
- Mentor junior engineers and help shape engineering culture as an early contributor.
Requirements
- 5+ years of hands-on experience designing, architecting, and scaling at least one production HPC or GPU compute cluster.
- Deep knowledge of server hardware, including GPUs, NICs, PCIe, memory, thermals, and power.
- Experience debugging performance and reliability issues across hardware, operating systems, drivers, and networking layers.
- Strong Linux systems administration experience, including kernel drivers, RDMA tuning, and performance analysis.
- Ability and willingness to mentor junior engineers and contribute to team culture.
- Willingness to work from the San Francisco office 3–4 days per week and travel domestically when required.
Nice to have
- Data center operations experience, including power, cooling, and colocation or vendor engagements.
- Experience with Slurm and Kubernetes.
- Exposure to KVM, QEMU, and libvirt.
- Experience with telemetry pipelines for predictive hardware failure detection.
- Experience troubleshooting InfiniBand or RoCEv2 Ethernet fabrics.
Culture & Benefits
- Competitive salary with company equity.
- Visa and work permit sponsorship.
- 401(k) matching up to 4%.
- Medical, dental, and vision insurance for employees and dependents, with 100% of premiums covered.
- Unlimited paid time off, 10+ observed holidays, and paid parental leave.
- Daily lunch and an office book budget.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →