обновлено 13 дней назад
Site Reliability Engineer (AI)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Making vLLM-powered inference services dependable in production with an accent on SLOs, observability, incident response, and operational automation. Focus on debugging distributed systems, improving deployment safety and recovery, and planning capacity for latency-sensitive, high-throughput AI workloads.
Location: San Francisco, California; remote work may be considered within the US for exceptional candidates
Salary: $200,000–$400,000 annual salary plus equity
Company
develops vLLM-powered AI inference infrastructure to make model inference faster and more cost-efficient.
What you will do
- Own reliability across the production lifecycle, including release safety, observability, incident response, and recovery.
- Define SLOs and SLIs for availability, latency, and successful inference requests.
- Build monitoring, dashboards, alerting, and automation for distributed inference services.
- Drive incident mitigation, escalation, root-cause analysis, and post-mortems with concrete prevention actions.
- Partner with engineering on capacity planning, operational readiness, and failure-mode reduction.
Requirements
- Bachelor’s degree or equivalent experience in computer science, engineering, systems, infrastructure, or a related field.
- Hands-on experience operating production distributed systems, cloud infrastructure, or reliability-critical services.
- Strong programming or scripting skills in Python, Go, Bash, or a similar language.
- Experience responding to production incidents, including mitigation, escalation, root-cause analysis, and corrective actions.
- Practical knowledge of SLOs, SLIs, error budgets, alerting, Linux, networking, and systems debugging.
- Clear communication and sound judgment when working with engineering teams under pressure.
Nice to have
- Experience with AI inference, model serving, ML infrastructure, GPU workloads, or latency-sensitive high-throughput services.
- Experience operating Kubernetes services and using Docker, Terraform, or comparable infrastructure tooling.
- Experience building observability with metrics, logs, traces, dashboards, and actionable alerts.
- Experience with deployment automation, release gates, rollback procedures, resource scheduling, or capacity planning.
Culture & Benefits
- Hands-on collaboration with engineers building the inference platform.
- Health, dental, and vision benefits.
- 401(k) company match.
- Equity included in the compensation package.
- Visa sponsorship is available on a case-by-case basis.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (AWS)
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
13 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
4 дня назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
152 000 - 228 000$
P2P.org
4 дня назад
Site Reliability Engineer (Web3)
5 000 - 6 000$
6 дней назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
132 000 - 211 400$