1 день назад
Site Reliability Engineer (AI Inference)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI Inference): Making vLLM-powered AI inference systems reliable, observable, and operationally simple at production scale with an accent on SLOs, incident response, and distributed systems. Focus on designing failure-resistant infrastructure, improving observability and recovery, and reducing operational risk for high-throughput inference workloads.
Location: San Francisco, California; remote work may be considered within the US for exceptional candidates
Salary: $200,000–$400,000 USD annually plus equity
Company
develops vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient.
What you will do
- Define SLOs, improve monitoring and alerting, and strengthen incident response for production inference systems.
- Lead mitigation, root cause analysis, escalation, and follow-up prevention work during major incidents.
- Drive post-mortems, reliability reviews, and operational readiness improvements.
- Design operationally simple systems and identify failure modes before launch.
- Improve service design, release safety, capacity planning, and production reliability with engineering teams.
- Build automation and tooling that reduce toil and improve recovery time.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, systems, infrastructure, or a related field.
- Strong experience operating production systems with significant traffic, user impact, or infrastructure criticality.
- Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.
- Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.
- Programming or scripting ability in Python, Go, Bash, or similar technologies.
- Experience with production incident mitigation, root cause analysis, escalation, and prevention work.
Nice to have
- Experience with ML infrastructure, AI inference systems, GPU workloads, Kubernetes platforms, or high-scale backend services.
- Experience with metrics, logs, traces, dashboards, alerts, and runbooks.
- Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD, or production deployment platforms.
- Experience owning reliability for high-throughput, latency-sensitive, or mission-critical systems.
- Experience leading severe outage response and communicating across engineering and leadership.
Culture & Benefits
- Work at the intersection of AI models and hardware on systems powering inference at scale.
- Generous health, dental, and vision benefits.
- 401(k) company match.
- Equity included in the compensation package.
- Visa sponsorship is available on a case-by-case basis.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Reliability Engineer (AI Infrastructure)
240 000 - 290 000$
1 день назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
2 дня назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
2 дня назад
Site Reliability Engineer (Kubernetes)
140 000 - 170 000$
1 день назад
Production Engineer (AI)
155 000 - 185 000$
2 дня назад
Senior Software Engineer (AWS/Kubernetes)
160 000 - 215 000$