Назад
Company hidden
14 часов назад

Senior Site Reliability Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI): Operating and scaling production systems for a healthcare AI workflow automation platform with an accent on reliability standards, observability, incident response, and distributed-system performance. Focus on designing SLOs and error budgets, reducing MTTR, optimizing AWS serverless and containerized services, and building automation that improves operational resilience.

Location: Hybrid in San Francisco, California; R&D roles require two days per week in the San Francisco office.

Company

hirify.global builds an AI workflow automation platform for healthcare operators across hospitals, health systems, pharmacies, and payors.

What you will do

  • Define and implement SLIs, SLOs, and error budgets for core services while owning uptime, latency, and availability targets.
  • Operate production systems, improve on-call rotations, and lead incident triage, mitigation, resolution, and blameless postmortems.
  • Design observability across metrics, logs, and distributed tracing using OpenTelemetry, Datadog, CloudWatch, Grafana, and Sentry.
  • Analyze performance under load and optimize latency, throughput, and resource usage across AWS Lambda, ECS, Aurora Postgres, and ClickHouse.
  • Build automation, incident-response tooling, deployment safeguards, CI/CD reliability checks, and capacity-planning tools.
  • Partner with security and compliance teams on operational standards, audit readiness, and monitoring integration with SIEM workflows.

Requirements

  • 5+ years of experience in Site Reliability Engineering, production infrastructure, or related roles.
  • Hands-on experience operating and debugging distributed systems in production.
  • Experience with observability tooling, incident response, on-call practices, and performance and reliability debugging.
  • Experience defining and using SLOs, SLIs, and error budgets.
  • Familiarity with AWS, serverless and container-based architectures, and Postgres or similar relational databases.
  • Ability to write Python, Bash, or similar scripts for automation and tooling.

Nice to have

  • Experience in high-growth or high-scale environments or regulated industries such as healthcare or fintech.
  • Experience with ClickHouse, analytical systems at scale, chaos engineering, or load testing.
  • Exposure to ML infrastructure or data platforms.

Culture & Benefits

  • Mission-driven team focused on improving healthcare operations and patient outcomes.
  • Flexible remote-first environment with meaningful office presence in San Francisco and New York.
  • Medical, dental, and vision insurance with family participation.
  • 401(k) plan with a company match and equity for full-time employees.
  • Unlimited PTO, paid parental leave, wellness and commuter stipends, and a weekly lunch stipend.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →