Senior Software Engineer (AI Cloud Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Software Engineer (AI Cloud Infrastructure): Build control-plane systems for Lambda’s GPU cloud, turning physical GPU infrastructure into reliable customer-facing capacity with an accent on distributed orchestration, schedulers, and operational tooling. Focus on designing robust APIs, workflow engines, and state-machine-driven lifecycle automation while improving reliability, observability, and production readiness.
Location: Hybrid — must be present in the San Francisco/San Jose office 4 days per week; work from home day is currently Tuesday.
Salary: $296K–$346K (annual)
Company
Lambda builds AI cloud infrastructure to make compute as ubiquitous as electricity.
What you will do
- Build and operate core cloud platform services for compute lifecycle, bare metal hosts, capacity, placement, and maintenance workflows.
- Design reliable backend services, APIs, state machines, and orchestration systems for GPU cloud control-plane operations.
- Own bare metal lifecycle workflows including launch, terminate, restart/reboot, host reclaim, validation, quarantine, and return-to-pool.
- Improve deployment readiness through observability, testing, alerting, runbooks, and operational tooling for business-critical services.
- Debug complex production issues across distributed services, infrastructure dependencies, networking, and cloud workflows.
- Partner across infrastructure, networking, fleet, security, support, and product teams to define cross-system contracts and deliver end-to-end capabilities.
Requirements
- 6+ years of professional software engineering experience building production backend or distributed systems.
- Strong in Python, Go, or a similar backend/system language.
- Experience designing and operating APIs, workflow engines, schedulers, orchestration services, or other distributed systems.
- Strong reliability fundamentals: fault tolerance, idempotency, retries, state machines, failure handling, and production debugging.
- Comfort with Linux, containers, Kubernetes, infrastructure automation, and service deployment patterns.
- Experience owning production services and participating in on-call, with improvements driven by operational learnings.
Culture & Benefits
- Hybrid schedule with required in-office presence 4 days per week; designated work-from-home day is Tuesday.
- Health, dental, and vision coverage for employees and dependents.
- 401k plan with 2% company match for USA employees.
- Flexible paid time off plan that employees actually use.
- Generous cash and equity compensation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →