1 день назад
Site Reliability Engineer (AI Infrastructure)
240 000 - 356 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Infrastructure): Build and operate monitoring, automate deployment and lifecycle of large-scale HPC clusters for AI workloads with an accent on cluster health, fabric, GPU, and job-level signals. Focus on troubleshooting complex cluster issues, automating remediation, and improving operational efficiency in a hybrid office environment.
Location
Location: Must be present in San Francisco or Bellevue office 4 days per week; designated work from home day is Tuesday
Salary: $240K – $356K per year
Company
Lambda is a leader in AI cloud infrastructure with 500+ employees, serving AI researchers, enterprises, and hyperscalers. Founded in 2012, the company focuses on making compute as ubiquitous as electricity.
What you will do
- Build and operate monitoring and alerting for cluster health including fabric, GPU, power/thermal, and job-level signals
- Deploy and configure large-scale HPC clusters for AI workloads using automation tools
- Automate cluster lifecycle management with Ansible and Terraform
- Create runbooks and automated remediations for common cluster failures
- Troubleshoot cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power
- Participate in on-call rotations and lead incident response for cluster problems
- Maintain Standard Operating Procedures and collaborate with engineering teams for operational improvements
Requirements
- 7+ years experience in Site Reliability Engineering, HPC Engineering, or DevOps
- Strong understanding of AI infrastructure, GPU architectures, and hardware performance optimization
- Experience with Linux-based distributed systems
- Proficiency in configuring and troubleshooting InfiniBand, RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL
- Solid knowledge of Python and Go, and experience improving internal tooling
- Experience with monitoring tools like Prometheus, Grafana, Clickhouse
- Proficiency in automation/configuration management tools such as Ansible and Terraform
Nice to have
- Experience with ML/DL frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
- Knowledge of containerization and orchestration (Docker, Kubernetes)
- Experience building or operating HPC resources
- Depth in NVIDIA hardware and firmware ecosystem
- Experience with data center power and thermal design
- Background in chaos engineering or reliability testing
- Understanding of compliance frameworks (SOC 2, ISO 27001)
Culture & Benefits
- Generous cash and equity compensation
- Health, dental, and vision coverage for employees and dependents
- Wellness and commuter stipends for select roles
- 401k plan with 2% company match (USA employees)
- Flexible paid time off plan
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Infrastructure Engineer (AI Hardware)
150 000 - 250 000$
3 дня назад
HPC Operations Engineer
175 000 - 225 000$
1 день назад
Senior Network Engineer (AI Infrastructure)
150 000 - 190 000$
3 дня назад
Cloud Support Engineer
145 000 - 175 000$
3 дня назад
Linux Device Management Engineer
160 000 - 200 000$
1 день назад
Infrastructure Engineer (Storage)
180 000 - 220 000$