Назад
5 дней назад

Senior or Staff Software Engineer (AI)

255 000 - 346 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior or Staff Software Engineer (AI): Building and operating distributed systems, APIs, control planes, and platform capabilities that power Lambda's GPU cloud with an accent on reliability, security, scalability, and operational excellence. Focus on solving complex state, consistency, concurrency, scheduling, and failure-recovery problems across customer-facing cloud infrastructure.

Location: Hybrid; presence required 4 days per week at the San Francisco, San Jose, or Bellevue office. Work from home day is Tuesday.

Annual salary: $255,000–$346,000 in San Francisco/San Jose; $230,000–$311,000 in Bellevue.

Company

Lambda builds AI cloud infrastructure for AI researchers, enterprises, and hyperscalers, with a focus on GPU-powered computing.

What you will do

  • Design, build, and operate services, APIs, control planes, and platform capabilities for the AI cloud.
  • Solve distributed-systems challenges involving state, consistency, concurrency, scheduling, failure recovery, and lifecycle management.
  • Own the full engineering lifecycle, including architecture, implementation, testing, rollout, observability, on-call, and continuous improvement.
  • Improve availability, latency, throughput, efficiency, security, and operability at large scale.
  • Turn incidents and near misses into automation, testing, guardrails, and other durable engineering improvements.
  • Collaborate with product, infrastructure, networking, storage, security, and SRE teams; contribute to reviews, standards, mentorship, and cross-team architecture at Staff level.

Requirements

  • 7+ years of professional software engineering experience or equivalent evidence of impact building production systems.
  • Strong proficiency in at least one general-purpose programming language; Go and Python are used primarily.
  • Experience designing, building, and operating backend services, distributed systems, infrastructure, or platform capabilities at meaningful scale.
  • Practical understanding of system design, data models, APIs, failure modes, performance, and production reliability tradeoffs.
  • Experience owning complex work through delivery and operation, including testing, staged rollout, monitoring, incident response, and root-cause improvement.
  • Ability to align cross-functional partners and build consensus around technical decisions and tradeoffs.

Nice to have

  • Experience with AWS, GCP, Azure, Kubernetes, container orchestration, schedulers, controllers, or cloud control-plane systems.
  • Experience in cloud infrastructure domains such as compute, storage, networking, identity and access, developer platforms, billing, databases, or fleet management.
  • Experience with infrastructure automation, durable workflows, event-driven architectures, or infrastructure as code.
  • Experience designing highly available, multi-region, or rapidly scaling distributed systems.
  • Familiarity with GPU infrastructure, HPC environments, or large-scale AI/ML training and inference workloads.

Culture & Benefits

  • Cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401(k) plan with a 2% company match for USA employees.
  • Flexible paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →