Назад
Company hidden
5 часов назад

Staff Site Reliability Engineer (AI/ML)

241 000 - 270 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (AI/ML): Architecting and operating reliable AWS and Kubernetes infrastructure for healthcare products and AI/ML workloads with an accent on SLOs, observability, incident response, and infrastructure automation. Focus on designing reliability standards, solving complex scaling and resilience challenges, optimizing cloud performance and cost, and building hands-free operational workflows with AI tools.

Location: Remote, with occasional travel to Garner's headquarters in New York City

Salary: $241,000–$270,000 per year, plus equity incentive and benefits

Company

hirify.global uses clinical metrics, healthcare data, and AI to help employers guide members toward higher-quality care while reducing healthcare costs.

What you will do

  • Own the reliability, performance, and resilience strategy for AWS and Kubernetes environments supporting healthcare products and AI/ML workloads.
  • Design the organization-wide SLO framework and lead technical direction for production quality and scalability.
  • Lead complex incident response, on-call improvements, root-cause analysis, and blameless post-incident reviews.
  • Architect monitoring, alerting, and observability platforms that detect and resolve issues proactively.
  • Translate ambiguous infrastructure needs into automated Terraform deliverables while improving cloud cost efficiency and performance.
  • Establish deployment and observability standards, mentor engineers, and ensure infrastructure changes meet security and HIPAA compliance requirements.

Requirements

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in SRE, DevOps, or platform engineering.
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment; AWS experience is preferred.
  • Experience designing SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews.
  • Strong Python or Go skills for infrastructure automation.
  • Experience driving cloud cost and performance optimization, setting technical direction, and mentoring engineers.
  • Excellent communication skills and experience applying AI tools to engineering and operations workflows.

Nice to have

  • Experience supporting AI/ML or data-intensive workloads in production.
  • Experience in security-conscious or regulated environments, including HIPAA or SOC 2.
  • Experience with the Kubernetes API.

Culture & Benefits

  • Remote work with occasional travel to the New York City headquarters.
  • Flexible paid time off.
  • Medical, dental, and vision plan options.
  • 401(k), Teladoc Health, equity incentives, and additional competitive benefits.
  • High-accountability environment focused on urgency, authentic feedback, and improving healthcare outcomes.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →