1 день назад
Site Reliability Engineer (AI Infrastructure)
175 000 - 265 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Infrastructure): Building and operating reliable infrastructure across colocation facilities, on-premises GPU clusters, cloud environments, and customer-facing platform services with an accent on automation, observability, and bare-metal operations. Focus on designing infrastructure as code, resolving incidents across the hardware-to-application stack, and supporting AI and HPC workloads with Kubernetes, cloud platforms, and high-performance interconnects.
Location: Santa Clara, United States; hybrid work
Salary: $175,000–$265,000 per year, plus equity and bonus opportunities
Company
develops software and hardware infrastructure for generative AI and focuses on collaborative, execution-oriented engineering.
What you will do
- Own reliability and availability across colocation server fleets, on-premises GPU clusters, cloud environments, and customer-facing platform services.
- Provision and operate infrastructure from bare metal through Kubernetes, including operating systems, networking, storage, hardware, and autoscaling.
- Automate provisioning, deployment, operations, host lifecycle management, fleet health checks, remediation, and networking using Terraform and/or Ansible.
- Design monitoring, alerting, dashboards, SLIs, and AIOps-driven detection workflows with tools such as Prometheus, Grafana, DataDog, and Splunk.
- Participate in on-call rotation, resolve incidents from bare metal to application layers, and produce root-cause analyses for P0/P1 incidents.
- Support CI/CD, QA, HPC workloads, operational runbooks, capacity planning, hardware lifecycle management, and cloud cost optimization.
Requirements
- 7+ years of experience in SRE, infrastructure engineering, or systems administration.
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- Strong Linux expertise with colocation or on-premises infrastructure, including networking, storage, systemd, kernel parameters, performance diagnostics, rack networking, and bare-metal provisioning.
- Production experience writing and maintaining Terraform and/or Ansible configurations.
- Operational Kubernetes experience covering cluster troubleshooting, workloads, storage, and networking.
- Experience with observability tools, production-quality Python and/or Bash scripting, incident response, structured triage, and RCA follow-through.
Nice to have
- Experience with customer-facing infrastructure and external reliability commitments.
- Cloud operations across AWS, Azure, or GCP in hybrid cloud and on-premises environments.
- Production experience with AIOps, intelligent alerting, anomaly detection, or LLM-assisted diagnostics.
- Experience with Slurm, LSF, InfiniBand, RoCE, NVLink, or large-scale infrastructure automation.
Culture & Benefits
- Collaborative and inclusive environment built around respect, humility, direct communication, and diverse perspectives.
- Medical, dental, and vision coverage.
- 401(k) and an employee rewards program focused on personal and family wellbeing.
- Equity and performance-based bonus opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Baseten
3 дня назад
Site Reliability Engineer (AI)
165 000 - 330 000$
5 дней назад
Site Reliability Engineer I (Azure)
6 дней назад
Sr. Site Reliability Engineer (AI)
1 день назад
Senior Site Reliability Engineer (AI)
75 000 - 85 000$
5 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
Microsoft AI
5 дней назад
AI Reliability Engineer (HPC)
119 800 - 234 700$