Назад
обновлено 8 дней назад

Principal Software Engineer (GPU Compute)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Principal Software Engineer (GPU Compute): Serving as the technical anchor for GPU and AI accelerator capabilities with an accent on machine management, reliability, and performance at scale. Focus on architecting GPU host lifecycles, optimizing driver/firmware stacks, and defining automated repair strategies for large-scale production environments.

Location: Must be based in or able to work from the San Mateo, CA office (Hybrid: Tue-Thu onsite)

Company

Roblox is a global platform building immersive 3D digital experiences that connect millions of people through shared play and creation.

What you will do

  • Lead GPU strategy by partnering across Kubernetes, Networking, and Cloud teams.
  • Own the GPU host lifecycle, including driver/firmware management, telemetry, and remediation of hardware faults.
  • Architect how GPU capacity is exposed to compute platforms via scheduling and isolation.
  • Drive GPU reliability and performance at fleet scale through automated diagnosis and repair.
  • Evaluate and onboard new GPU/AI accelerator platforms and networking topologies.
  • Establish standards and APIs to enable other engineering teams to consume GPU compute efficiently.

Requirements

  • 10+ years of experience in large-scale distributed systems and infrastructure.
  • Deep, hands-on expertise in GPU host provisioning, driver lifecycles, and production reliability.
  • Strong proficiency in Go or other well-structured programming languages.
  • Experience operating GPU/AI workloads with CUDA, GPU scheduling, and high-performance networking.
  • Must have existing United States work authorization (no H-1B sponsorship support).

Nice to have

  • Familiarity with Kubernetes for GPU workloads.
  • Experience with bare-metal concepts like firmware, BMC/IPMI/Redfish, and OS imaging.

Culture & Benefits

  • Hybrid work model with onsite presence required Tuesday through Thursday.
  • Comprehensive benefits package including health, equity compensation, and more.
  • Opportunity to solve unique technical challenges at massive scale.
  • Commitment to equal employment opportunity and inclusive work environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →