Назад
3 дня назад

Senior Software Engineer (AI)

255 000 - 346 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer (AI/Kubernetes): Building managed Kubernetes, orchestration, and inference platform services for large-scale AI training and inference with an accent on GPU-aware scheduling, distributed systems, and cloud-native infrastructure. Focus on designing resilient control planes, integrating NVIDIA networking and GPU technologies, and automating cluster and model-serving lifecycles.

Location: Hybrid; presence required 4 days per week in the San Francisco, San Jose, or Bellevue office. Tuesday is the designated work-from-home day.

Annual salary: $255,000–$346,000 in San Francisco/San Jose; $230,000–$311,000 in Bellevue.

Company

Lambda builds AI cloud infrastructure for researchers, enterprises, hyperscalers, and large-scale machine learning workloads.

What you will do

  • Design and maintain scalable control-plane services, operators, and custom Kubernetes controllers.
  • Develop Go and Python automation for cluster provisioning, upgrades, patching, and deletion.
  • Build GPU-aware orchestration for scheduling and resource allocation across AI workloads.
  • Integrate networking technologies including Cilium, Multus, InfiniBand, RoCE, RDMA, and GPUDirect.
  • Develop inference platform services, autoscaling systems, multi-model deployment patterns, and internal deployment and monitoring tools.
  • Support production systems through debugging and an on-call rotation.

Requirements

  • 6+ years of software engineering experience and ownership of significant technical scope.
  • Deep knowledge of Kubernetes internals, including controllers, schedulers, operators, CRDs, CSI, CNI, and control-plane components.
  • Strong understanding of distributed systems, fault tolerance, graceful degradation, and failure handling.
  • Strong programming skills in Go and Python, with solid knowledge of Linux, networking, containers, and cloud infrastructure.
  • Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting.
  • Ability to work from the required United States office location on the hybrid schedule.

Nice to have

  • Experience building managed Kubernetes services such as GKE, EKS, or AKS.
  • Hands-on experience with NVIDIA GPU and networking technologies, including GPU Operator, device plugins, DCGM, MIG, Network Operator, and NCCL.
  • Familiarity with Slurm, KAI, Volcano, Kueue, InfiniBand, RDMA, high-performance computing, and AI/ML storage architecture.
  • Contributions to CNCF projects or Kubernetes SIGs.

Culture & Benefits

  • Work on core infrastructure used by AI research labs, enterprises, and hyperscalers.
  • Collaborate across Kubernetes, networking, storage, compute, ML, and infrastructure teams.
  • Generous cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles, a 401(k) plan with a 2% company match for US employees, and flexible paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →