Назад
15 часов назад

Staff Software Engineer (AI Cloud)

314 000 - 419 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Software Engineer (AI Cloud): Building a highly available GPU and CPU host and instance lifecycle control plane for Lambda's large-scale AI cloud with an accent on distributed systems, hardware enablement, and cloud provisioning. Focus on designing resilient compute infrastructure, integrating BIOS/firmware and DPU capabilities, and leading multi-team delivery of secure, enterprise-grade systems.

Location: Hybrid, with presence in the Bellevue, San Francisco, or San Jose office 4 days per week; Tuesday is the designated work-from-home day.

Salary: $314,000–$419,000 annually in Bellevue or $349,000–$465,000 annually in San Francisco/San Jose.

Company

Lambda provides AI cloud infrastructure for researchers, enterprises, and hyperscalers, focused on making large-scale compute broadly accessible.

What you will do

  • Design and implement a highly available GPU and CPU host and instance lifecycle control plane.
  • Define technical approaches spanning semiconductor architecture, BIOS/UEFI, firmware, boot methodologies, and DPU utilization.
  • Design secure multi-tenant compute platform capabilities.
  • Lead technical direction, mentorship, design reviews, and engineering standards across multiple teams.
  • Collaborate with product, data center organizations, and customers to turn technical requirements into scalable infrastructure deliverables.

Requirements

  • 10+ years of experience building resilient distributed systems for deploying and managing heterogeneous compute platforms in data centers.
  • Deep expertise in durable execution models and distributed systems for cloud-service provisioning.
  • Experience leading semiconductor hardware enablement and deploying new data centers into a global compute platform.
  • Knowledge of software-defined networking fundamentals and secure multi-tenant systems.
  • Proficiency in one or more of C, C++, Rust, Python, or Go.

Nice to have

  • Knowledge of NVIDIA AI Factory components and software, including GPU/CPU hosts, SuperNICs, ConnectX, BlueField DPUs, DOCA, and CUDA.
  • Knowledge of Linux kernel internals, device drivers, KVM, QEMU, SR-IOV, DPDK, or SPDK.
  • Experience with cloud-provider Kubernetes offerings, InfiniBand, RoCE, or NVMe-oF.

Culture & Benefits

  • Cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for USA employees.
  • Flexible paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →