Назад
1 день назад

Site Reliability Engineer (AI Infrastructure)

240 000 - 356 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI Infrastructure): Build and operate monitoring, automate deployment and lifecycle of large-scale HPC clusters for AI workloads with an accent on cluster health, fabric, GPU, and job-level signals. Focus on troubleshooting complex cluster issues, automating remediation, and improving operational efficiency in a hybrid office environment.

Location

Location: Must be present in San Francisco or Bellevue office 4 days per week; designated work from home day is Tuesday

Salary: $240K – $356K per year

Company

Lambda is a leader in AI cloud infrastructure with 500+ employees, serving AI researchers, enterprises, and hyperscalers. Founded in 2012, the company focuses on making compute as ubiquitous as electricity.

What you will do

  • Build and operate monitoring and alerting for cluster health including fabric, GPU, power/thermal, and job-level signals
  • Deploy and configure large-scale HPC clusters for AI workloads using automation tools
  • Automate cluster lifecycle management with Ansible and Terraform
  • Create runbooks and automated remediations for common cluster failures
  • Troubleshoot cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power
  • Participate in on-call rotations and lead incident response for cluster problems
  • Maintain Standard Operating Procedures and collaborate with engineering teams for operational improvements

Requirements

  • 7+ years experience in Site Reliability Engineering, HPC Engineering, or DevOps
  • Strong understanding of AI infrastructure, GPU architectures, and hardware performance optimization
  • Experience with Linux-based distributed systems
  • Proficiency in configuring and troubleshooting InfiniBand, RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL
  • Solid knowledge of Python and Go, and experience improving internal tooling
  • Experience with monitoring tools like Prometheus, Grafana, Clickhouse
  • Proficiency in automation/configuration management tools such as Ansible and Terraform

Nice to have

  • Experience with ML/DL frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
  • Knowledge of containerization and orchestration (Docker, Kubernetes)
  • Experience building or operating HPC resources
  • Depth in NVIDIA hardware and firmware ecosystem
  • Experience with data center power and thermal design
  • Background in chaos engineering or reliability testing
  • Understanding of compliance frameworks (SOC 2, ISO 27001)

Culture & Benefits

  • Generous cash and equity compensation
  • Health, dental, and vision coverage for employees and dependents
  • Wellness and commuter stipends for select roles
  • 401k plan with 2% company match (USA employees)
  • Flexible paid time off plan

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →