Назад
Company hidden
обновлено 2 дня назад

Senior Site Reliability Engineer (AI)

191 000 - 226 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AWS/Kubernetes/Terraform): Building and operating reliable, resilient cloud infrastructure for healthcare products and AI/ML workloads with an accent on SLOs, observability, incident response, and infrastructure automation. Focus on scaling production systems, eliminating operational toil with AI tools, optimizing cloud performance and cost, and meeting HIPAA security requirements.

Location: Remote, with occasional travel to the New York City headquarters

Salary: $191,000–$226,000 per year, plus equity and benefits

Company

hirify.global uses clinical metrics, healthcare data, and incentives to help employers and members access higher-quality, lower-cost care in the United States.

What you will do

  • Own the reliability, performance, and resilience of AWS and Kubernetes cloud environments, including infrastructure supporting AI/ML workloads.
  • Define and uphold SLOs, participate in on-call rotation, lead incident response, and drive root-cause analysis and corrective actions.
  • Build and maintain monitoring, alerting, and observability systems.
  • Translate scaling requirements into automated, composable Terraform infrastructure and improve cloud cost efficiency and performance.
  • Use AI tools and automation to eliminate operational toil, reduce technical debt, and create monitored, hands-free processes.
  • Establish deployment, observability, security, and compliance standards while supporting engineering teams and communicating with technical and non-technical stakeholders.

Requirements

  • 4+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS experience is preferred.
  • Experience defining SLOs, building monitoring and alerting, leading incident response, and conducting blameless post-incident reviews.
  • Strong software engineering fundamentals in Python or Go, applied to infrastructure automation.
  • Experience with cloud cost and performance optimization across compute, storage, and networking.
  • Fluency with AI tools applied to engineering and operations workflows, or strong motivation to develop this capability quickly.

Nice to have

  • Experience supporting AI/ML or data-intensive production workloads.
  • Experience operating in security-conscious or regulated environments, including HIPAA or SOC 2.
  • Experience with Kubernetes APIs.

Culture & Benefits

  • Mission-driven work focused on improving healthcare outcomes for millions of people.
  • High-performance environment with individual accountability, urgency, and authentic feedback.
  • Flexible PTO and medical, dental, and vision plan options.
  • 401(k) with company match, flexible spending accounts, Teladoc Health, equity participation, and additional benefits.
  • Visa sponsorship is not available.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →