Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Engineer (Managed Kubernetes/AI): Building Lambda’s bare-metal managed Kubernetes platform and related orchestration services for GPU-accelerated AI training and inference with an accent on distributed systems, NVIDIA’s ecosystem, and full-stack infrastructure design. Focus on GPU-aware scheduling, multi-tenancy, high availability, managed Slurm, inference platforms, self-healing automation, and chaos engineering at massive scale.
Location: Hybrid, with presence required four days per week in the San Francisco, San Jose, or Bellevue office; Tuesday is the designated work-from-home day.
Annual salary: $349,000–$465,000 in San Francisco/San Jose or $314,000–$419,000 in Bellevue.
Company
Lambda builds AI cloud infrastructure for researchers, enterprises, and hyperscalers, with a focus on GPU-accelerated computing and managed services.
What you will do
- Set the technical vision for a bare-metal managed Kubernetes platform, including control-plane scalability, multi-tenancy, cluster lifecycle management, and high availability.
- Build GPU-aware orchestration services by integrating NVIDIA GPU Operator, Network Operator, DCGM, NCCL, AICR, Topograph, and related tooling.
- Develop the foundation for Managed Slurm on Kubernetes and higher-level inference services, including model serving, autoscaling, and multi-model deployment.
- Define networking, storage, compute, and security requirements for AI workloads in collaboration with infrastructure teams.
- Design self-healing systems, incident-response automation, observability, chaos engineering, upgrade automation, security patching, and zero-downtime maintenance.
- Lead technical direction across the Orchestration team, mentor engineers, collaborate with customers and NVIDIA, and contribute to the open-source community.
Requirements
- 10+ years of software, platform engineering, or SRE experience, including at least 5 years focused on Kubernetes at scale.
- Expert knowledge of Kubernetes internals, including API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
- Strong production software engineering skills in Go and Python.
- Deep experience with GPU orchestration, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.
- Expertise across distributed systems, compute, networking, storage, security, Linux, observability, infrastructure-as-code, and GitOps.
- Proven technical leadership experience with managed services or multi-tenant platforms, including design decisions, mentoring, and cross-team influence.
Nice to have
- Experience with NVIDIA Network Operator, GPUDirect, NCCL tuning, Topograph, AICR, or similar projects.
- Experience with Slurm, KAI, Volcano, Kueue, managed Kubernetes services, or Kubernetes control-plane components.
- Background in confidential computing, ML infrastructure, customer migrations, security and compliance, or CNCF, Kubernetes SIG, and NVIDIA open-source contributions.
Culture & Benefits
- Work on foundational AI cloud infrastructure used by research labs, enterprises, and large AI companies.
- Direct collaboration with NVIDIA and engineers specializing in machine learning, systems, and infrastructure.
- Cash and equity compensation, health, dental, and vision coverage, and flexible paid time off.
- 401(k) plan with a 2% company match for USA employees.
- Wellness and commuter stipends for select roles.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Staff Software Engineer (Kubernetes)
215 000 - 265 000$
5 дней назад
Infrastructure Engineer (AI)
250 000 - 300 000$
6 дней назад
Staff Software Engineer, Cloud Infrastructure (AI)
181 000 - 265 000$
Adyen
5 дней назад
Senior Platform Engineer (Kubernetes)
180 000 - 243 000$
6 дней назад
Platform Engineer (AI)
250 000 - 315 000$
8 дней назад
Software Engineer, Platform Operations (Kubernetes)
127 000 - 158 700$