Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Backend Engineer (AI): Developing and maintaining production systems that power GPU fleet lifecycle management and machine configuration at scale with an accent on automation frameworks for machine provisioning and system health monitoring. Focus on debugging hardware and firmware issues and collaborating across teams to develop scalable and maintainable solutions.
Location: This position requires presence in our San Francisco/San Jose or Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.
Salary: San Francisco / San Jose $225K – $300K; Bellevue $203K – $270K
Company
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers.
What you will do
- Design, implement, and improve software that powers GPU fleet lifecycle management and machine configuration at scale.
- Build and enhance automation frameworks for machine provisioning, configuration management, and deployment.
- Enable bring-up, validation, and production readiness for new server and accelerator platforms.
- Improve and refine workflows for bare metal provisioning, firmware updates, and system health monitoring.
- Investigate failures across BIOS, BMC, firmware, networking, storage, and boot flows.
- Collaborate with infrastructure, security, and product engineering teams to develop scalable and maintainable solutions.
Requirements
- 2+ years of experience working with Go (Golang) or Python in production environments.
- 2+ years of experience with configuration management tools and practices.
- Comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers.
- Can independently troubleshoot complex systems and communicate effectively across software, infrastructure, and vendor teams.
Nice to have
- Experience with Go in infrastructure, systems, or backend development.
- Hands-on experience with bare metal provisioning and lifecycle management, including technologies such as Redfish, BMC, IPMI, DHCP, and PXE.
- Experience diagnosing issues involving drivers, firmware, and hardware compatibility across GPU servers.
- Experience incorporating AI-assisted development tools into engineering workflows, including code generation, debugging, test development, and documentation.
- Experience building Linux distributions or managing OS customization and imaging.
- Familiarity with Ansible for system configuration and automation.
- Exposure to Kubernetes and container orchestration concepts.
Culture & Benefits
- Generous cash & equity compensation.
- Health, dental, and vision coverage for you and your dependents.
- Wellness and commuter stipends for select roles.
- 401k Plan with 2% company match (USA employees).
- Flexible paid time off plan that we all actually use.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Backend Engineer (AI)
150 000 - 300 000$
6 дней назад
Senior Software Engineer (AI)
190 000 - 210 000$
4 дня назад
Senior Backend Engineer (AI)
180 000 - 250 000$
Snowflake
6 дней назад
Backend Infrastructure Engineer (AI)
200 000 - 270 000$
5 дней назад
Senior Backend Engineer (AI)
130 000 - 162 200$
Snowflake
5 дней назад
Staff Software Engineer (AI-first)
160 000 - 230 000$