9 часов назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI): Designing and operating highly available, scalable cloud-native systems for AI products with an accent on production reliability, observability, and distributed systems. Focus on incident response, defining SLOs and error budgets, building reliability automation, and supporting multi-region or GPU-intensive platforms.
Location: San Francisco, California, United States. Hybrid work with a minimum of three in-office days per week.
Company
builds AI hardware and software products that capture, structure, and utilize intelligence from conversations for more than 2,000,000 users worldwide.
What you will do
- Ensure the reliability and performance of .ai AI products at scale.
- Design and operate highly available, scalable cloud-native systems for AI workloads.
- Own production reliability, on-call practices, and incident response.
- Build observability through metrics, logs, tracing, and reliability automation.
- Define and manage SLOs, SLIs, and error budgets with engineering teams.
- Lead postmortems and continuous reliability improvements across the platform.
Requirements
- 5+ years of experience in SRE, infrastructure, or platform engineering.
- Strong experience with AWS, GCP, Azure, or OCI.
- Hands-on experience with Kubernetes and distributed systems.
- Experience with on-call rotations and incident management.
- Proficiency in at least one programming language: Go, Python, or Java.
Nice to have
- Experience supporting AI/ML or data-intensive platforms.
- GPU cluster management experience.
- Knowledge of SLO/SLA frameworks.
- Experience with fast-growing or global products and multi-region systems.
- Strong written and verbal communication skills.
Culture & Benefits
- Employee Stock Ownership Plan with a stake in the company’s long-term success.
- Medical, dental, and vision insurance with employer support, plus a 401(k) matching plan.
- Unlimited PTO, 13 paid holidays, and 12 weeks of fully paid parental leave.
- Access to AI productivity tools, high-performance equipment, and devices.
- Annual offsites, team events, and a product-driven culture focused on craftsmanship, ownership, and velocity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
15 часов назад
Senior Site Reliability Engineer (AI)
7 часов назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
12 часов назад
Platform Engineer (AI)
16 часов назад
Senior Platform Engineer (AI)
9 часов назад
DevOps Engineer (AI)
183 000 - 248 000$
8 часов назад
Senior Site Reliability Engineer (AWS)
180 000 - 200 000$