17 часов назад
Staff SRE (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff SRE (AI): Building self-service reliability platforms and GitOps-driven delivery systems for high-performance AI inference infrastructure across datacenters and cloud environments with an accent on observability, automation, and SLO-driven operations. Focus on architecting model release pipelines, capacity provisioning, cluster upgrades, and reliability practices for latency, throughput, and accuracy at scale.
Location: Hybrid in the SF Bay Area or Toronto
Company
builds wafer-scale AI hardware and high-performance inference infrastructure for model labs, enterprises, and AI-native startups.
What you will do
- Define and implement strategies for reliable software delivery and operations across multiple datacenters and cloud environments.
- Architect self-service platforms and internal tooling for product teams, external customers, and cluster operators.
- Build GitOps-driven delivery for model releases, capacity provisioning, and cluster upgrades.
- Define SLOs, SLIs, error budgets, blameless postmortems, chaos testing, and capacity forecasting for inference workloads.
- Mentor SREs, support critical incident escalations, and prioritize automation based on production pain points.
- Measure toil reduction, deployment velocity, SLO compliance, MTTR, and self-service adoption.
Requirements
- 8+ years of experience in SRE, infrastructure engineering, or platform engineering.
- Experience improving automation and reliability at large scale in demanding technical environments.
- Deep expertise operating large heterogeneous clusters with a proprietary cloud control plane.
- Experience designing and delivering CI/CD or GitOps systems with Argo CD or similar tools.
- Hands-on experience with Loki, Tempo, Mimir, Prometheus, or comparable observability systems.
- Ability to lead complex projects, influence stakeholders, and communicate technical direction clearly.
Nice to have
- Production experience with Bazel or other large-scale build systems.
- Experience with AI/ML inference, model serving runtimes, GPU or wafer-scale orchestration, and latency or accuracy SLOs.
- Experience with predictive autoscaling, chaos engineering, or cost-aware capacity planning.
Culture & Benefits
- Work on a breakthrough AI platform and high-performance AI supercomputer.
- Opportunities to publish and open-source AI research.
- Startup vitality combined with job stability.
- Simple, non-corporate work culture that respects individual beliefs.
- No 24/7 on-call rotation is required.
- Commitment to an inclusive and diverse work environment with continuous learning and support.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
19 часов назад
Staff Production Engineer (AI)
209 000 - 253 000$
Bitdeer
6 дней назад
Senior SRE Platform Architect (AI)
23 часа назад
Site Reliability Engineer (Observability)
3 дня назад
Senior Software Engineer (SRE)
Windsurf
7 дней назад
Site Reliability Engineer (AI)
23 часа назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$