Назад
Company hidden
1 день назад

Site Reliability Engineer (AI Inference)

200 000 - 400 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI Inference): Making vLLM-powered AI inference systems reliable, observable, and operationally simple at production scale with an accent on SLOs, incident response, and distributed systems. Focus on designing failure-resistant infrastructure, improving observability and recovery, and reducing operational risk for high-throughput inference workloads.

Location: San Francisco, California; remote work may be considered within the US for exceptional candidates

Salary: $200,000–$400,000 USD annually plus equity

Company

hirify.global develops vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient.

What you will do

  • Define SLOs, improve monitoring and alerting, and strengthen incident response for production inference systems.
  • Lead mitigation, root cause analysis, escalation, and follow-up prevention work during major incidents.
  • Drive post-mortems, reliability reviews, and operational readiness improvements.
  • Design operationally simple systems and identify failure modes before launch.
  • Improve service design, release safety, capacity planning, and production reliability with engineering teams.
  • Build automation and tooling that reduce toil and improve recovery time.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, systems, infrastructure, or a related field.
  • Strong experience operating production systems with significant traffic, user impact, or infrastructure criticality.
  • Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.
  • Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.
  • Programming or scripting ability in Python, Go, Bash, or similar technologies.
  • Experience with production incident mitigation, root cause analysis, escalation, and prevention work.

Nice to have

  • Experience with ML infrastructure, AI inference systems, GPU workloads, Kubernetes platforms, or high-scale backend services.
  • Experience with metrics, logs, traces, dashboards, alerts, and runbooks.
  • Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD, or production deployment platforms.
  • Experience owning reliability for high-throughput, latency-sensitive, or mission-critical systems.
  • Experience leading severe outage response and communicating across engineering and leadership.

Culture & Benefits

  • Work at the intersection of AI models and hardware on systems powering inference at scale.
  • Generous health, dental, and vision benefits.
  • 401(k) company match.
  • Equity included in the compensation package.
  • Visa sponsorship is available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →