Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Software Engineer, HPC (distributed compute infrastructure): Building core services, APIs, scheduling systems, and orchestration strategies for a high-scale HPC platform supporting thousands of concurrent jobs with an accent on reliability, resource efficiency, and multi-region performance. Focus on designing distributed systems, optimizing heterogeneous workload scheduling, and building developer platforms for large engineering organizations.
Location: Hybrid in Foster City, California; Boston, Massachusetts; or Seattle, Washington
Company
Develops autonomous mobility products supported by large-scale software, compute, and infrastructure platforms.
What you will do
- Design and implement distributed compute services and abstractions for thousands of concurrent workloads.
- Build production-grade APIs, SDKs, and developer tools for large-scale distributed computing.
- Develop scheduling algorithms, autoscaling policies, and multi-region orchestration strategies.
- Investigate systemic reliability and performance issues through profiling, analysis, and collaboration with workload owners.
- Create capacity planning and forecasting tools for growing compute requirements.
- Lead cross-team infrastructure initiatives, define the HPC roadmap, and mentor junior engineers.
Requirements
- Experience designing and operating large-scale distributed systems in production.
- Experience with Ray.io, especially Ray Core and Ray Data, or equivalent technologies.
- Experience using Kubernetes for heterogeneous workloads and AWS or similar cloud infrastructure.
- Track record of delivering and operating reliable, scalable infrastructure.
- Proficiency with Python.
- Ability to prioritize engineering work and build cross-functional consensus around technical trade-offs.
Nice to have
- Experience with machine learning workloads, including training, inference, or data generation.
- Experience operating Kubernetes or SLURM at scale exceeding 10,000 nodes.
- Background in SLURM workload management, advanced scheduling policies, algorithmic optimization, or operations research.
- Experience building developer tools and platforms for large engineering organizations.
Culture & Benefits
- Full-time employment in a hybrid work environment.
- Cross-functional collaboration with infrastructure and customer-facing engineering teams.
- Opportunity to lead multi-quarter initiatives with organization-wide impact.
- Mentorship and support for junior engineers' career development.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Senior Backend Software Engineer (HPC)
10 дней назад
Senior/Staff Backend Engineer (Geothermal)
10 дней назад
Senior/Staff Software Engineer
12 дней назад
Software Engineer (Java or Python) (Cybersecurity)
8 дней назад
Staff Software Engineer (Python)
10 дней назад
Senior Software Engineer, Backend (Platform) (Python/AWS)
180 000 - 220 000$