Назад
3 дня назад

Staff Engineer (Managed Kubernetes)

349 000 - 465 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Engineer (Managed Kubernetes/AI): Building Lambda’s bare-metal managed Kubernetes platform and related orchestration services for GPU-accelerated AI training and inference with an accent on distributed systems, NVIDIA’s ecosystem, and full-stack infrastructure design. Focus on GPU-aware scheduling, multi-tenancy, high availability, managed Slurm, inference platforms, self-healing automation, and chaos engineering at massive scale.

Location: Hybrid, with presence required four days per week in the San Francisco, San Jose, or Bellevue office; Tuesday is the designated work-from-home day.

Annual salary: $349,000–$465,000 in San Francisco/San Jose or $314,000–$419,000 in Bellevue.

Company

Lambda builds AI cloud infrastructure for researchers, enterprises, and hyperscalers, with a focus on GPU-accelerated computing and managed services.

What you will do

  • Set the technical vision for a bare-metal managed Kubernetes platform, including control-plane scalability, multi-tenancy, cluster lifecycle management, and high availability.
  • Build GPU-aware orchestration services by integrating NVIDIA GPU Operator, Network Operator, DCGM, NCCL, AICR, Topograph, and related tooling.
  • Develop the foundation for Managed Slurm on Kubernetes and higher-level inference services, including model serving, autoscaling, and multi-model deployment.
  • Define networking, storage, compute, and security requirements for AI workloads in collaboration with infrastructure teams.
  • Design self-healing systems, incident-response automation, observability, chaos engineering, upgrade automation, security patching, and zero-downtime maintenance.
  • Lead technical direction across the Orchestration team, mentor engineers, collaborate with customers and NVIDIA, and contribute to the open-source community.

Requirements

  • 10+ years of software, platform engineering, or SRE experience, including at least 5 years focused on Kubernetes at scale.
  • Expert knowledge of Kubernetes internals, including API machinery, controllers, schedulers, operators, CRDs, CSI, CNI, and extension patterns.
  • Strong production software engineering skills in Go and Python.
  • Deep experience with GPU orchestration, including NVIDIA GPU Operator, device plugins, DCGM, MIG, time-slicing, and GPU-aware scheduling.
  • Expertise across distributed systems, compute, networking, storage, security, Linux, observability, infrastructure-as-code, and GitOps.
  • Proven technical leadership experience with managed services or multi-tenant platforms, including design decisions, mentoring, and cross-team influence.

Nice to have

  • Experience with NVIDIA Network Operator, GPUDirect, NCCL tuning, Topograph, AICR, or similar projects.
  • Experience with Slurm, KAI, Volcano, Kueue, managed Kubernetes services, or Kubernetes control-plane components.
  • Background in confidential computing, ML infrastructure, customer migrations, security and compliance, or CNCF, Kubernetes SIG, and NVIDIA open-source contributions.

Culture & Benefits

  • Work on foundational AI cloud infrastructure used by research labs, enterprises, and large AI companies.
  • Direct collaboration with NVIDIA and engineers specializing in machine learning, systems, and infrastructure.
  • Cash and equity compensation, health, dental, and vision coverage, and flexible paid time off.
  • 401(k) plan with a 2% company match for USA employees.
  • Wellness and commuter stipends for select roles.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →