7 часов назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Ensuring the reliability and performance of Plaud’s AI products by designing scalable cloud-native systems and operating production infrastructure with an accent on Kubernetes, distributed systems, and observability. Focus on incident response, SLO and error-budget management, reliability automation, and supporting AI workloads at scale.
Location: Seattle, WA; hybrid with a minimum of three in-office days per week
Company
builds hardware-software AI products that capture, structure, and use intelligence from conversations and other real-world interactions.
What you will do
- Ensure the reliability and performance of ’s AI products at scale.
- Design and operate highly available, scalable, cloud-native systems for AI workloads.
- Own production reliability, incident response, on-call practices, and postmortems.
- Build observability across metrics, logs, tracing, and reliability automation.
- Define and manage SLOs, SLIs, and error budgets with engineering teams.
- Partner with product and engineering teams to improve reliability design and operational maturity.
Requirements
- 5+ years of experience in SRE, infrastructure, or platform engineering roles.
- Strong experience with at least one major cloud platform: AWS, GCP, Azure, or OCI.
- Hands-on experience with Kubernetes and distributed systems.
- Experience with on-call rotations and incident management.
- Proficiency in at least one programming language: Go, Python, or Java.
Nice to have
- Experience supporting AI/ML or data-intensive platforms.
- GPU cluster management experience.
- Knowledge of SLO/SLA frameworks and multi-region systems.
- Experience with fast-growing or global products.
- Strong written and verbal communication skills.
Culture & Benefits
- Employee Stock Ownership Plan with a stake in ’s long-term success.
- Medical, dental, and vision insurance for employees and dependents, plus a company-matched 401(k) plan.
- Unlimited PTO, 13 paid holidays, and 12 weeks of fully paid parental leave.
- Access to AI productivity tools, high-performance equipment, and devices.
- Annual company offsites, team events, and a culture focused on craftsmanship, ownership, innovation, and velocity.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 часов назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
7 часов назад
Infrastructure Engineer (AI)
8 часов назад
Infrastructure Engineer (AI)
158 000 - 235 000$
5 дней назад
Infrastructure Engineer (AI)
200 000 - 350 000$
15 часов назад
Senior Site Reliability Engineer (AI)
9 часов назад
Infrastructure Engineer (AI)
200 000 - 400 000$