5 часов назад
Staff Site Reliability Engineer (AI/ML)
241 000 - 270 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (AI/ML): Architecting and operating reliable AWS and Kubernetes infrastructure for healthcare products and AI/ML workloads with an accent on SLOs, observability, incident response, and infrastructure automation. Focus on designing reliability standards, solving complex scaling and resilience challenges, optimizing cloud performance and cost, and building hands-free operational workflows with AI tools.
Location: Remote, with occasional travel to Garner's headquarters in New York City
Salary: $241,000–$270,000 per year, plus equity incentive and benefits
Company
uses clinical metrics, healthcare data, and AI to help employers guide members toward higher-quality care while reducing healthcare costs.
What you will do
- Own the reliability, performance, and resilience strategy for AWS and Kubernetes environments supporting healthcare products and AI/ML workloads.
- Design the organization-wide SLO framework and lead technical direction for production quality and scalability.
- Lead complex incident response, on-call improvements, root-cause analysis, and blameless post-incident reviews.
- Architect monitoring, alerting, and observability platforms that detect and resolve issues proactively.
- Translate ambiguous infrastructure needs into automated Terraform deliverables while improving cloud cost efficiency and performance.
- Establish deployment and observability standards, mentor engineers, and ensure infrastructure changes meet security and HIPAA compliance requirements.
Requirements
- 7+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS experience is preferred.
- Experience designing SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews.
- Strong Python or Go skills for infrastructure automation.
- Experience driving cloud cost and performance optimization, setting technical direction, and mentoring engineers.
- Excellent communication skills and experience applying AI tools to engineering and operations workflows.
Nice to have
- Experience supporting AI/ML or data-intensive workloads in production.
- Experience in security-conscious or regulated environments, including HIPAA or SOC 2.
- Experience with the Kubernetes API.
Culture & Benefits
- Remote work with occasional travel to the New York City headquarters.
- Flexible paid time off.
- Medical, dental, and vision plan options.
- 401(k), Teladoc Health, equity incentives, and additional competitive benefits.
- High-accountability environment focused on urgency, authentic feedback, and improving healthcare outcomes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 часа назад
Senior DevOps Engineer / Site Reliability Engineer (AI)
170 000 - 220 000$
7 часов назад
Staff Software Engineer (AWS/Kubernetes)
175 000 - 245 000$
SandboxAQ
6 дней назад
Staff Platform Engineer (AI)
121 600 - 228 000$
8 часов назад
Senior Platform Engineer (AI)
170 000 - 240 000$
11 часов назад
Principal Site Reliability Engineer (AWS/Kubernetes)
163 620 - 212 710$
5 дней назад
Senior Site Reliability Engineer (Kubernetes)
93 700 - 138 700$