обновлено 2 дня назад
Senior Site Reliability Engineer (AI)
191 000 - 226 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes/Terraform): Building and operating reliable, resilient cloud infrastructure for healthcare products and AI/ML workloads with an accent on SLOs, observability, incident response, and infrastructure automation. Focus on scaling production systems, eliminating operational toil with AI tools, optimizing cloud performance and cost, and meeting HIPAA security requirements.
Location: Remote, with occasional travel to the New York City headquarters
Salary: $191,000–$226,000 per year, plus equity and benefits
Company
uses clinical metrics, healthcare data, and incentives to help employers and members access higher-quality, lower-cost care in the United States.
What you will do
- Own the reliability, performance, and resilience of AWS and Kubernetes cloud environments, including infrastructure supporting AI/ML workloads.
- Define and uphold SLOs, participate in on-call rotation, lead incident response, and drive root-cause analysis and corrective actions.
- Build and maintain monitoring, alerting, and observability systems.
- Translate scaling requirements into automated, composable Terraform infrastructure and improve cloud cost efficiency and performance.
- Use AI tools and automation to eliminate operational toil, reduce technical debt, and create monitored, hands-free processes.
- Establish deployment, observability, security, and compliance standards while supporting engineering teams and communicating with technical and non-technical stakeholders.
Requirements
- 4+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering.
- Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS experience is preferred.
- Experience defining SLOs, building monitoring and alerting, leading incident response, and conducting blameless post-incident reviews.
- Strong software engineering fundamentals in Python or Go, applied to infrastructure automation.
- Experience with cloud cost and performance optimization across compute, storage, and networking.
- Fluency with AI tools applied to engineering and operations workflows, or strong motivation to develop this capability quickly.
Nice to have
- Experience supporting AI/ML or data-intensive production workloads.
- Experience operating in security-conscious or regulated environments, including HIPAA or SOC 2.
- Experience with Kubernetes APIs.
Culture & Benefits
- Mission-driven work focused on improving healthcare outcomes for millions of people.
- High-performance environment with individual accountability, urgency, and authentic feedback.
- Flexible PTO and medical, dental, and vision plan options.
- 401(k) with company match, flexible spending accounts, Teladoc Health, equity participation, and additional benefits.
- Visa sponsorship is not available.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Staff Platform Infrastructure Engineer (AI)
180 000 - 220 000$
Writer
2 дня назад
Infrastructure Engineer (AI)
155 400 - 273 700$
6 дней назад
Senior SecDevOps Engineer (AI)
165 000 - 200 000$
7 дней назад
Senior Machine Learning Platform Engineer (AI)
170 000 - 220 000$
6 дней назад
Staff Platform Engineer (AI/ML)
140 800 - 176 000$
7 дней назад
Staff Forward Deployed Platform Engineer (AI)
185 000 - 265 000$