Назад
5 дней назад

Senior / Principal Infrastructure Engineer - ML Platform (AI)

278 530 - 345 040$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior / Principal Infrastructure Engineer - ML Platform (AI): Building and scaling foundational Kubernetes and cloud infrastructure for Roblox's machine learning platform with an accent on model serving, GPU fleet management, and reliable high-scale systems. Focus on designing custom Kubernetes controllers, automating infrastructure with Terraform, orchestrating hybrid-cloud environments, and improving the end-to-end ML lifecycle for data scientists and ML engineers.

Location: San Mateo, California, United States; office-based roles require onsite attendance Tuesday through Thursday, with optional presence on Monday and Friday.

Salary: $278,530–$345,040 USD annual base pay.

Company

Roblox develops a platform for creating and experiencing 3D immersive digital experiences used by a global community of developers and creators.

What you will do

  • Bootstrap and maintain Kubernetes and cloud infrastructure for the ML Platform serving layer, metadata store, model registry, and pipeline orchestrator.
  • Define technical strategy and oversee the development of reliable, high-scale infrastructure systems.
  • Develop platform tooling that improves time to production across the machine learning lifecycle.
  • Build infrastructure for GPU fleet management, hybrid-cloud orchestration, and custom Kubernetes controllers and resources.
  • Partner with data scientists and ML engineers to create tooling, interfaces, and visualizations for ML platform users.

Requirements

  • 6+ years of professional engineering experience with system design for scalable, reliable platforms.
  • Deep experience managing Kubernetes clusters at scale, including hundreds to thousands of nodes and 100k+ QPS.
  • Strong Terraform and infrastructure-as-code experience across AWS, GCP, or similar cloud environments.
  • Proficiency with Docker, Kubernetes, CI/CD systems, and cloud infrastructure.
  • Experience with model serving, training, model CI/CD, and GPU resource management.
  • Bachelor's degree in Computer Science, Computer Engineering, Data Science, or a similar technical field, or equivalent practical experience.

Culture & Benefits

  • Work on infrastructure supporting hundreds of machine learning use cases and billions of inferences per day.
  • Collaborate across organizations to improve the experience of internal ML platform users.
  • Full-time employees are eligible for equity compensation and company benefits.
  • Equal employment opportunities and reasonable accommodations are provided during the recruiting process.
  • Future H-1B sponsorship may not be supported, and certain US visa categories may not be eligible for employment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →