Назад
1 месяц назад

Staff Software Engineer, HPC

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Software Engineer, HPC (distributed compute infrastructure): Building core services, APIs, scheduling systems, and orchestration strategies for a high-scale HPC platform supporting thousands of concurrent jobs with an accent on reliability, resource efficiency, and multi-region performance. Focus on designing distributed systems, optimizing heterogeneous workload scheduling, and building developer platforms for large engineering organizations.

Location: Hybrid in Foster City, California; Boston, Massachusetts; or Seattle, Washington

Company

Develops autonomous mobility products supported by large-scale software, compute, and infrastructure platforms.

What you will do

  • Design and implement distributed compute services and abstractions for thousands of concurrent workloads.
  • Build production-grade APIs, SDKs, and developer tools for large-scale distributed computing.
  • Develop scheduling algorithms, autoscaling policies, and multi-region orchestration strategies.
  • Investigate systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
  • Create capacity planning and forecasting tools for growing compute requirements.
  • Lead cross-team infrastructure initiatives, define the HPC roadmap, and mentor junior engineers.

Requirements

  • Experience designing and operating large-scale distributed systems in production.
  • Experience with Ray.io, especially Ray Core and Ray Data, or equivalent technologies.
  • Experience using Kubernetes for heterogeneous workloads and AWS or similar cloud infrastructure.
  • Track record of delivering and operating reliable, scalable infrastructure.
  • Proficiency with Python.
  • Ability to prioritize engineering work and build cross-functional consensus around technical trade-offs.

Nice to have

  • Experience with machine learning workloads, including training, inference, or data generation.
  • Experience operating Kubernetes or SLURM at scale exceeding 10,000 nodes.
  • Background in SLURM workload management, advanced scheduling policies, algorithmic optimization, or operations research.
  • Experience building developer tools and platforms for large engineering organizations.

Culture & Benefits

  • Full-time employment in a hybrid work environment.
  • Cross-functional collaboration with infrastructure and customer-facing engineering teams.
  • Opportunity to lead multi-quarter initiatives with organization-wide impact.
  • Mentorship and support for junior engineers' career development.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →