5 дней назад
SRE (AI Inference Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE (AI Inference Infrastructure) (Kubernetes/Python/Go): Operating and scaling production infrastructure for a high-performance AI inference service with an accent on releases, capacity changes, cluster upgrades, and observability. Focus on building self-service continuous delivery pipelines, reusable automation, and reliability practices for frontier-class AI workloads.
Location: San Francisco Bay Area or Toronto, United States/Canada
Company
builds high-performance AI hardware and infrastructure for training and inference workloads.
What you will do
- Execute production releases, capacity changes, and cluster upgrades while the SRE function scales.
- Develop self-service continuous delivery pipelines using Kubernetes, Bazel, Prometheus, Grafana, InfluxDB, Python, and Go.
- Build reusable automation and internal developer tools to reduce operational toil and cross-team friction.
- Develop telemetry, observability, and alerting solutions for reliable operations at scale.
- Collaborate with Cluster Ops and development teams to identify and implement high-impact automation.
- Contribute to SLOs, post-mortems, reliability metrics, and capacity planning.
Requirements
- 2–4+ years of SRE experience with a strong operations or automation focus.
- Production experience with Kubernetes.
- Proficiency in Python or Go for tools and automation.
- Experience with Prometheus, Grafana, and observability-driven workflows.
- Ability to measure and communicate reliability, operational toil, and velocity impact.
Nice to have
- Hands-on GitOps experience with Argo CD, Flux, or an equivalent tool.
- Experience building continuous delivery pipelines.
- Experience with Bazel or similar build systems.
- Familiarity with capacity planning and on-premises or multi-datacenter environments.
Culture & Benefits
- Work on a high-performance AI platform and wafer-scale computing architecture.
- Opportunity to work with cutting-edge AI infrastructure and leading model builders.
- Direct mentorship from experienced engineers and ownership of production systems.
- No requirement for 24/7 on-call rotations.
- Startup vitality, job stability, and a simple, non-corporate work culture.
- Continuous learning, growth, and support in an inclusive workplace.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Infrastructure Engineer (AI)
5 дней назад
Production Engineer (AI Infrastructure)
172 000 - 209 000$
5 дней назад
Site Reliability Engineer (Observability)
5 дней назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
5 дней назад
Senior Production Engineer (AI Infrastructure)
170 000 - 205 000$
5 дней назад