3 часа назад
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract (AI)
195 000 - 285 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Director, Site Reliability Engineering - AI Accelerator Infrastructure - Contract (AI): Building and leading the SRE function for colocation facilities, on-premises lab clusters, multi-cloud environments, and customer-facing AI accelerator infrastructure with an accent on reliability architecture, observability, automation, and capacity planning. Focus on establishing SLOs and incident management, designing shared storage and hybrid-cloud platforms, and scaling infrastructure operations across HPC workloads and silicon programs.
Location: Santa Clara, United States; hybrid
Salary: $195,000–$285,000 per year, plus equity and bonus opportunities
Company
designs and manufactures purpose-built AI inference silicon and develops the supporting software, hardware, QA, research, and infrastructure systems.
What you will do
- Build and lead the SRE function, including its charter, technical roadmap, team structure, and operational standards.
- Hire, develop, and retain a team of 3–5 SRE engineers while directing data center and lab technician operations.
- Establish SLOs, error budgets, on-call rotations, incident management, observability, and RCA practices from a zero baseline.
- Own reliability across colocation, on-premises lab clusters, AWS, Azure, GCP, and customer-facing platform services.
- Drive infrastructure-as-code, self-healing automation, capacity planning, FinOps, and migration to enterprise shared storage.
- Partner with DevOps, hardware, software, and executive stakeholders on infrastructure reliability and HPC workload requirements.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field, plus 15+ years of SRE, infrastructure, or production engineering experience.
- 5+ years leading SRE or infrastructure engineering teams, including building or significantly rebuilding an SRE function.
- Deep Linux, TCP/IP, RDMA, bare-metal, enterprise storage, colocation, on-premises hardware, and Kubernetes experience.
- Production-scale Terraform and Ansible expertise, including module design, remote state, environment isolation, and change governance.
- Experience owning Prometheus, Grafana, and/or Datadog observability and writing production services or automation in Python and/or Go.
- Strong executive communication and the ability to create structure in a high-ambiguity, low-process environment.
Nice to have
- Experience with customer-facing infrastructure, InfiniBand, RoCE, NVLink, or HPC schedulers such as Slurm and LSF.
- Multi-cloud hybrid operations, FinOps, ITIL, technical writing, conference talks, or open-source contributions.
Culture & Benefits
- Collaborative, inclusive culture centered on respect, humility, direct communication, and execution.
- Medical, dental, and vision coverage.
- 401(k) and an inclusive rewards plan supporting employee wellbeing and dependents.
- Six-month contract with the possibility of conversion to a full-time role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Director of AI Infrastructure (AI)
176 400 - 264 600$
3 часа назад
Senior Engineering Manager (AI Infrastructure)
146 880 - 220 320$
Nscale
5 дней назад
Software Engineering Manager (AI Infrastructure)
300 000 - 350 000$
3 часа назад
Head of Customer Engineering (AI)
315 000 - 375 000$
5 дней назад
Staff Engineering Manager (AI)
140 400 - 372 300$
Scale AI
2 дня назад
Sr. Director, Forward Deployed Engineering (AI)
285 200 - 356 500$