Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Building reliable multi-cloud Kubernetes infrastructure, observability tooling, and automated mitigations for an ML infrastructure platform with an accent on SLOs, incident response, and runtime performance. Focus on diagnosing latency, memory, GPU utilization, concurrency, and model lifecycle issues while developing self-healing systems and operational runbooks.
Location: Hybrid in San Francisco, Montreal, New York, or Toronto
Salary: $165K–$330K annually, plus equity
Company
Baseten provides AI inference infrastructure, applied AI research, and developer tooling for deploying machine learning models in production.
What you will do
- Own the reliability of multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.
- Build and maintain observability infrastructure for metrics, logging, dashboards, and alerting as code.
- Define and instrument SLOs and SLIs across customer workloads and internal services.
- Create runbooks and convert recurring failure patterns into automated mitigations and self-healing systems.
- Diagnose runtime issues involving latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.
- Collaborate with engineering, forward-deployed, and product teams to improve operational practices.
Requirements
- Extensive hands-on Kubernetes experience and experience building scalable infrastructure.
- Strong observability foundation with metrics, logging, dashboards, and alerting pipelines.
- Experience with infrastructure as code using Terraform or Helm and GitOps workflows.
- Experience writing runbooks, leading incident response, and conducting post-mortem analysis.
- Ability to combine software engineering with operational process design, escalation planning, and incident management.
- Curiosity about deploying and serving machine learning models at scale; prior ML experience is not required.
Nice to have
- Multi-cloud Kubernetes experience with EKS, GKE, or similar platforms.
- Observability-as-code experience.
- Familiarity with incident.io or a similar incident management platform.
Culture & Benefits
- Competitive compensation with meaningful equity.
- Flexible PTO and a company-wide Winter Break.
- Paid parental leave and a fertility and family-building stipend.
- U.S.-only: Medical, dental, and vision insurance coverage for employees and dependents.
- U.S.-only: Company-facilitated 401(k).
- Exposure to AI startups and production machine learning systems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
6 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$
4 дня назад
Head of Site Reliability Engineering (AI)
195 000 - 285 000$
8 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
9 дней назад
Site Reliability Engineer
90 000 - 110 000$
Replit
5 дней назад
Engineering Manager (SRE)
250 000 - 325 000$