5 дней назад
Principal SRE (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal SRE (AI): Architecting self-service delivery, observability, capacity orchestration, rollout safety, and operational automation for large-scale AI inference infrastructure with an accent on reliability, fleet management, and production control planes. Focus on building unified capacity management systems, designing SLO-based reliability practices, and enabling safe self-service workflows across datacenters and cloud environments.
Location: On-site in the SF Bay Area or Toronto
Company
Systems develops wafer-scale AI hardware and high-performance inference infrastructure for model labs, enterprises, and AI-native startups.
What you will do
- Define and implement strategies for reliably delivering and operating software across multiple datacenters and cloud environments.
- Architect self-service platforms and internal tooling for safely triggering and observing critical workflows.
- Design reliability practices for inference workloads, including SLOs, SLIs, error budgets, postmortems, chaos testing, and capacity forecasting.
- Support incident escalations, mentor senior SREs, and prioritize automation based on production pain points.
- Measure impact through toil reduction, deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
Requirements
- 15+ years of experience in SRE, infrastructure engineering, or platform engineering in large-scale production environments.
- Deep experience with compute fleets, control planes, schedulers, orchestration systems, capacity management, and reliability automation.
- Experience driving cross-team architecture for production control planes, fleet management, capacity orchestration, or self-service infrastructure platforms.
- Ability to lead complex technical programs, influence senior stakeholders, mentor engineers, and communicate technical strategy.
- Hands-on experience with observability, incident response, and SLO-based reliability management across metrics, logs, traces, alerting, and dashboards.
Nice to have
- Experience with Bazel or other large-scale build systems.
- Background in AI/ML inference systems, model serving, disaggregated inference, GPU orchestration, latency and accuracy SLOs, or drift monitoring.
- Experience with predictive autoscaling, chaos engineering, or cost-aware capacity management.
Culture & Benefits
- Work on a wafer-scale AI platform designed beyond traditional GPU constraints.
- Opportunities to publish and open-source AI research.
- Access to one of the fastest AI supercomputers in the world.
- Startup vitality combined with job stability.
- Non-corporate culture focused on individual beliefs, learning, growth, and support.
- The role does not require 24/7 on-call rotations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
4 дня назад
SRE L1 Support/Cloud Platform Ops Engineer (AI)
5 дней назад
Senior Production Engineer (AI Infrastructure)
170 000 - 205 000$
5 дней назад
Staff Production Engineer (AI)
209 000 - 253 000$
CrowdStrike
3 дня назад
Sr Engineer, SRE TechOps CICD (Remote)
140 000 - 215 000$
3 дня назад