14 часов назад
Senior Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI): Operating and scaling production systems for a healthcare AI workflow automation platform with an accent on reliability standards, observability, incident response, and distributed-system performance. Focus on designing SLOs and error budgets, reducing MTTR, optimizing AWS serverless and containerized services, and building automation that improves operational resilience.
Location: Hybrid in San Francisco, California; R&D roles require two days per week in the San Francisco office.
Company
builds an AI workflow automation platform for healthcare operators across hospitals, health systems, pharmacies, and payors.
What you will do
- Define and implement SLIs, SLOs, and error budgets for core services while owning uptime, latency, and availability targets.
- Operate production systems, improve on-call rotations, and lead incident triage, mitigation, resolution, and blameless postmortems.
- Design observability across metrics, logs, and distributed tracing using OpenTelemetry, Datadog, CloudWatch, Grafana, and Sentry.
- Analyze performance under load and optimize latency, throughput, and resource usage across AWS Lambda, ECS, Aurora Postgres, and ClickHouse.
- Build automation, incident-response tooling, deployment safeguards, CI/CD reliability checks, and capacity-planning tools.
- Partner with security and compliance teams on operational standards, audit readiness, and monitoring integration with SIEM workflows.
Requirements
- 5+ years of experience in Site Reliability Engineering, production infrastructure, or related roles.
- Hands-on experience operating and debugging distributed systems in production.
- Experience with observability tooling, incident response, on-call practices, and performance and reliability debugging.
- Experience defining and using SLOs, SLIs, and error budgets.
- Familiarity with AWS, serverless and container-based architectures, and Postgres or similar relational databases.
- Ability to write Python, Bash, or similar scripts for automation and tooling.
Nice to have
- Experience in high-growth or high-scale environments or regulated industries such as healthcare or fintech.
- Experience with ClickHouse, analytical systems at scale, chaos engineering, or load testing.
- Exposure to ML infrastructure or data platforms.
Culture & Benefits
- Mission-driven team focused on improving healthcare operations and patient outcomes.
- Flexible remote-first environment with meaningful office presence in San Francisco and New York.
- Medical, dental, and vision insurance with family participation.
- 401(k) plan with a company match and equity for full-time employees.
- Unlimited PTO, paid parental leave, wellness and commuter stipends, and a weekly lunch stipend.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
15 часов назад
Senior Platform Engineer (AI)
7 часов назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
9 часов назад
Senior Site Reliability Engineer (AI)
9 часов назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
7 часов назад
Senior SRE Engineer (AI)
6 часов назад