3 дня назад
Senior Reliability Engineer (AI)
140 000 - 165 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Reliability Engineer (AI): Building observability, incident-response automation, and resilience tooling for critical healthcare journeys with an accent on SLO engineering, peak-load readiness, and AI-assisted operational excellence. Focus on designing incident pipelines, debugging failures across frontend, API, mesh, and database layers, and validating systems under high traffic.
Location: US Remote; compensation range applies to US-based candidates
Salary: $140K–$165K annually, plus potential equity
Company
Hims & Hers operates a health and wellness platform providing personalized access to diagnosis, treatment, and delivery.
What you will do
- Own reliability for Tier 1 customer journeys such as checkout, telehealth visits, and prescription fulfillment by defining SLOs, golden signals, and business-level monitors.
- Lead capacity planning, load testing, performance analysis, and resilience improvements for high-traffic events.
- Build connected incident-response workflows across FireHydrant, Datadog, Jira, and Confluence, including alerting, RCA creation, runbooks, and post-mortem tracking.
- Develop reusable AI agents and operational tooling for incident triage, RCA drafting, report generation, and observability analysis.
- Deep-debug cross-service failures across frontend, APIs, service mesh, and databases, then turn findings into runbooks and engineering improvements.
- Coach engineers on SLOs, blameless post-mortems, on-call practices, and operational excellence.
Requirements
- 5+ years of experience as a Software, SRE, Platform, or Infrastructure Engineer owning reliability outcomes for production systems.
- Strong software engineering skills and the ability to investigate application code across the stack.
- Hands-on experience with observability, SLOs, golden signals, burn-rate alerting, and actionable monitoring.
- Production experience with AWS, Kubernetes/EKS, Terraform, and PostgreSQL, including RDS/Aurora.
- Experience with incident management, on-call design, escalation policies, incident command, blameless post-mortems, and action-item follow-through.
- Practical use of AI coding and analysis tools, plus experience building AI agents or LLM-backed operational automation.
Nice to have
- Load-testing and performance-engineering experience with k6 or similar tools.
- Familiarity with Istio service mesh observability and traffic management.
- Healthcare or other regulated-environment experience.
- Experience designing vendor and partner escalation frameworks with severities and response SLAs.
Culture & Benefits
- Competitive salary and equity compensation for full-time roles.
- Unlimited PTO, company holidays, and quarterly mental health days.
- Medical, dental, vision, and parental-leave benefits.
- Employee Stock Purchase Program and 401(k) with employer matching.
- Offsite team retreats and a flexible remote work approach.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
Site Reliability Engineer (AI Security)
137 300 - 205 900$
7 дней назад
Sr. Manager, Site Reliability
59 550 - 110 594GBP
5 дней назад
Senior Engineer Site Reliability (AWS/Data Operations)
105 700 - 149 275$
10 дней назад
Senior Site Reliability Engineer (AWS/AI)
140 000 - 160 000$
3 дня назад
Site Reliability Engineer
100 000 - 130 000$
Datadog
9 дней назад
Senior Software Engineer - Incident Insights & Readiness (SRE)
192 000 - 240 000$