Назад
3 дня назад

Site Reliability Engineer (AI)

165 000 - 330 000$
Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (AI): Building reliable multi-cloud Kubernetes infrastructure, observability tooling, and automated mitigations for an ML infrastructure platform with an accent on SLOs, incident response, and runtime performance. Focus on diagnosing latency, memory, GPU utilization, concurrency, and model lifecycle issues while developing self-healing systems and operational runbooks.

Location: Hybrid in San Francisco, Montreal, New York, or Toronto

Salary: $165K–$330K annually, plus equity

Company

Baseten provides AI inference infrastructure, applied AI research, and developer tooling for deploying machine learning models in production.

What you will do

  • Own the reliability of multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking.
  • Build and maintain observability infrastructure for metrics, logging, dashboards, and alerting as code.
  • Define and instrument SLOs and SLIs across customer workloads and internal services.
  • Create runbooks and convert recurring failure patterns into automated mitigations and self-healing systems.
  • Diagnose runtime issues involving latency, memory behavior, GPU utilization, concurrency, and model lifecycle management.
  • Collaborate with engineering, forward-deployed, and product teams to improve operational practices.

Requirements

  • Extensive hands-on Kubernetes experience and experience building scalable infrastructure.
  • Strong observability foundation with metrics, logging, dashboards, and alerting pipelines.
  • Experience with infrastructure as code using Terraform or Helm and GitOps workflows.
  • Experience writing runbooks, leading incident response, and conducting post-mortem analysis.
  • Ability to combine software engineering with operational process design, escalation planning, and incident management.
  • Curiosity about deploying and serving machine learning models at scale; prior ML experience is not required.

Nice to have

  • Multi-cloud Kubernetes experience with EKS, GKE, or similar platforms.
  • Observability-as-code experience.
  • Familiarity with incident.io or a similar incident management platform.

Culture & Benefits

  • Competitive compensation with meaningful equity.
  • Flexible PTO and a company-wide Winter Break.
  • Paid parental leave and a fertility and family-building stipend.
  • U.S.-only: Medical, dental, and vision insurance coverage for employees and dependents.
  • U.S.-only: Company-facilitated 401(k).
  • Exposure to AI startups and production machine learning systems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →