4 дня назад
Head of Site Reliability Engineering (AI)
195 000 - 285 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Head of Site Reliability Engineering (AI) (AI infrastructure): Building and leading the SRE function for AI inference silicon infrastructure across colocation facilities, on-premises labs, cloud environments, and customer-facing platform services with an accent on reliability engineering, observability, infrastructure automation, and storage architecture. Focus on defining SLOs, designing incident response and on-call systems, operating hybrid-cloud infrastructure, and leading the migration to enterprise-grade shared storage.
Location: Santa Clara, United States; hybrid
Salary: $195,000–$285,000 per year, plus equity and bonus opportunities
Company
develops purpose-built AI inference silicon and the software and infrastructure that support generative AI applications.
What you will do
- Build and lead the SRE function, define its charter, establish SRE practices, and hire and develop a team of 1–3 SRE engineers.
- Direct Data Center & Lab Technician operations across on-premises and colocation facilities.
- Own 24×7 reliability for colocation, on-premises lab clusters, cloud environments, and customer-facing platform services.
- Define SLIs, SLOs, error budgets, on-call rotations, incident management, and root-cause analysis processes.
- Own observability, infrastructure-as-code automation, self-healing systems, FinOps, capacity planning, and workload placement.
- Lead the migration from ad-hoc JBOD storage to an enterprise-grade shared storage platform across on-premises, colocation, and cloud environments.
Requirements
- 15+ years of experience in SRE, infrastructure engineering, or production engineering, plus 5+ years leading SRE or infrastructure engineering teams.
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
- Deep Linux expertise, including bare-metal operations, NAS/SAN, NFS/SMB, snapshots, replication, and hybrid-cloud storage.
- Production-scale Terraform and Ansible experience, Kubernetes operations, and hands-on operation of colocation and on-premises hardware.
- Experience with observability platforms such as Prometheus, Grafana, Datadog, or Splunk, plus strong Python and/or Go scripting skills.
- Ability to communicate infrastructure health and operational risk to senior and non-technical stakeholders and build structure in ambiguous environments.
Nice to have
- Experience with customer-facing infrastructure, multi-cloud hybrid operations across AWS, Azure, and GCP, or FinOps.
- Knowledge of InfiniBand, RoCE, NVLink, HPC schedulers such as Slurm or LSF, and structured incident or change management frameworks.
- Technical writing, conference presentations, or open-source contributions in reliability, observability, or HPC infrastructure.
Culture & Benefits
- Collaborative and inclusive environment focused on respect, humility, direct communication, and execution.
- Medical, dental, and vision coverage.
- 401(k) and an employee wellbeing-focused rewards package.
- Equity and performance-based bonus opportunities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Staff Site Reliability Engineer (Production Engineer) - Federal
119 000 - 170 000$
8 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
5 дней назад
Site Reliability Engineer Staff (Cloud Infrastructure)
6 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$
9 дней назад
Site Reliability Engineer
90 000 - 110 000$
6 дней назад
Principal Site Reliability Engineer (Kubernetes)
190 000 - 220 000$