3 дня назад
Infrastructure Engineer (GPU & Compute)
180 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Infrastructure Engineer (GPU & Compute) (AI infrastructure): Bringing up, validating, and operating large-scale bare-metal GPU compute infrastructure with an accent on system diagnostics, image pipelines, and automation. Focus on diagnosing hardware, OS, driver, and GPU issues, improving validation coverage, and building reliable provisioning workflows for AI/ML and HPC workloads.
Location: Hybrid role based in New York City, San Francisco, Seattle, or London, with a minimum of 2 in-office days per week; occasional team and company offsites.
Annual base salary: $180,000–$220,000 USD, plus discretionary bonus, equity, and benefits.
Company
builds an end-to-end platform for developing, training, and deploying AI systems, combining developer-focused software with large-scale compute infrastructure.
What you will do
- Own image management, deployment, provisioning, and validation systems for large-scale bare-metal infrastructure.
- Run test clusters and validate firmware, drivers, operating system images, and GPU-enabled systems.
- Develop GPU diagnostics and performance validation workflows using tools such as NVIDIA DCGM.
- Build Python-based automation to improve the reliability, repeatability, and scalability of infrastructure bring-up.
- Operate Linux production and validation environments, virtualization platforms, PXE provisioning, and image-based systems.
- Collaborate with Infrastructure, Hardware, Data Center, platform, and ML teams to qualify systems for AI/ML and HPC workloads.
Requirements
- 5+ years of experience in infrastructure engineering, systems engineering, or a related role.
- Strong Linux systems experience in production environments.
- Hands-on experience with GPU-enabled systems and NVIDIA DCGM.
- Experience with bare-metal provisioning, system bring-up, and debugging across hardware, operating systems, GPUs, and system software.
- Proficiency in Python or a similar scripting or programming language for automation.
- Visa sponsorship is not available for this position.
Nice to have
- Experience with InfiniBand, NVLink, PXE boot, LiveCD systems, or image-based provisioning.
- Experience with iDRAC, IPMI, Redfish, data center operations, or physical hardware.
- Experience supporting AI/ML or HPC workloads at scale.
- Experience with GPU validation frameworks or large-scale hardware qualification.
Culture & Benefits
- Flexible schedules and a hybrid work model for office-based teams.
- Medical, dental, and vision coverage, with eligible dependent coverage.
- RSUs, 401(k) matching in the U.S., and pension contributions in the U.K.
- Unlimited PTO, company holidays, floating holidays, and a two-week winter break.
- Paid parental and family leave, professional development allowance, wellness and work-from-home stipends.
- Four weeks of paid sabbatical leave after four years and complimentary meals at office hubs.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
5 дней назад
Systems Operations Support Engineer (AI Infrastructure)
90 000 - 160 000$
1 час назад
Hardware Systems Engineer (AI)
202 000 - 241 000$
5 дней назад
GPU Systems Engineer
200 000 - 300 000$
Lambda
3 дня назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$
5 дней назад
Systems Software Engineer (AI Cloud)
120 000 - 180 000$