Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Software Engineer (GPU Compute): Building reliable GPU and AI accelerator infrastructure for large-scale compute workloads with an accent on host lifecycle management, accelerator health, scheduling, and performance. Focus on designing automated detection and repair, integrating GPU capacity with Kubernetes, and supporting multi-node training and inference across data centers and cloud environments.
Location: Headquarters in San Mateo, California; office-based roles require onsite attendance Tuesday through Thursday, with optional presence on Monday and Friday.
Company
Roblox develops a platform for creating and experiencing 3D immersive digital experiences used by a global community of developers and creators.
What you will do
- Set the technical direction for GPU and AI accelerator capabilities across the Compute team.
- Own GPU host lifecycle management, including drivers, firmware, CUDA, health telemetry, and failure remediation.
- Architect GPU capacity exposure, scheduling, isolation, and Kubernetes integration for GPU and AI workloads.
- Improve GPU reliability and performance at fleet scale through detection, diagnosis, and automated repair.
- Evaluate new accelerators, networking topologies, and multi-node training and inference patterns.
- Establish tooling, standards, and APIs for safe and efficient GPU compute consumption across engineering teams.
Requirements
- 10+ years of experience building and operating large-scale distributed systems and infrastructure.
- Deep hands-on expertise in GPU host provisioning, driver and firmware lifecycle, GPU health, and production accelerator reliability.
- Experience operating GPU and AI workloads in production, including CUDA, GPU scheduling, and high-performance networking.
- Strong proficiency in Go or another well-structured programming language.
- Expertise with NVLink, InfiniBand, or RoCE and experience with multi-node workloads.
- Leadership experience as a technical anchor for complex GPU and compute problems.
Nice to have
- Experience with Kubernetes for GPU workloads.
- Familiarity with bare-metal systems, including firmware, BMC/IPMI/Redfish, and OS imaging.
Culture & Benefits
- Work on technical challenges involving GPU infrastructure at large scale.
- Collaborate across Kubernetes, Machine Bootstrap, Networking, and Cloud teams.
- Full-time employees are eligible for equity compensation and company benefits.
- Roblox provides reasonable accommodations during the recruiting process for qualifying disabilities or religious beliefs.
Hiring process
- Equal employment opportunity and non-discrimination principles apply throughout the recruiting process.
- For US-based roles, certain US visa categories may not be supported, and future H-1B sponsorship may not be available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
CoreWeave
11 дней назад
Senior Software Engineer (GPU & Runtime Systems)
153 000 - 204 000$
11 дней назад
Senior Software Engineer, Simulation System (Golang)
12 дней назад
Senior Software Engineer, Optimization Systems (Aerospace)
12 дней назад
Senior Software Engineer, Network On-Ramps (Networking)
12 дней назад
Software Engineer, Optimization Systems (Aerospace)
10 дней назад
Senior Backend Engineer (Distributed Systems)
215 000 - 265 000$