Назад
обновлено 24 дня назад

Principal Software Engineer (GPU Compute)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Software Engineer (GPU Compute): Building reliable GPU and AI accelerator infrastructure for large-scale compute workloads with an accent on host lifecycle management, accelerator health, scheduling, and performance. Focus on designing automated detection and repair, integrating GPU capacity with Kubernetes, and supporting multi-node training and inference across data centers and cloud environments.

Location: Headquarters in San Mateo, California; office-based roles require onsite attendance Tuesday through Thursday, with optional presence on Monday and Friday.

Company

Roblox develops a platform for creating and experiencing 3D immersive digital experiences used by a global community of developers and creators.

What you will do

  • Set the technical direction for GPU and AI accelerator capabilities across the Compute team.
  • Own GPU host lifecycle management, including drivers, firmware, CUDA, health telemetry, and failure remediation.
  • Architect GPU capacity exposure, scheduling, isolation, and Kubernetes integration for GPU and AI workloads.
  • Improve GPU reliability and performance at fleet scale through detection, diagnosis, and automated repair.
  • Evaluate new accelerators, networking topologies, and multi-node training and inference patterns.
  • Establish tooling, standards, and APIs for safe and efficient GPU compute consumption across engineering teams.

Requirements

  • 10+ years of experience building and operating large-scale distributed systems and infrastructure.
  • Deep hands-on expertise in GPU host provisioning, driver and firmware lifecycle, GPU health, and production accelerator reliability.
  • Experience operating GPU and AI workloads in production, including CUDA, GPU scheduling, and high-performance networking.
  • Strong proficiency in Go or another well-structured programming language.
  • Expertise with NVLink, InfiniBand, or RoCE and experience with multi-node workloads.
  • Leadership experience as a technical anchor for complex GPU and compute problems.

Nice to have

  • Experience with Kubernetes for GPU workloads.
  • Familiarity with bare-metal systems, including firmware, BMC/IPMI/Redfish, and OS imaging.

Culture & Benefits

  • Work on technical challenges involving GPU infrastructure at large scale.
  • Collaborate across Kubernetes, Machine Bootstrap, Networking, and Cloud teams.
  • Full-time employees are eligible for equity compensation and company benefits.
  • Roblox provides reasonable accommodations during the recruiting process for qualifying disabilities or religious beliefs.

Hiring process

  • Equal employment opportunity and non-discrimination principles apply throughout the recruiting process.
  • For US-based roles, certain US visa categories may not be supported, and future H-1B sponsorship may not be available.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →