Назад
Company hidden
15 часов назад

Director, Site Reliability Operations (AI-enabled Cloud Operations)

Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
director
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Director, Site Reliability Operations (AI-enabled Cloud Operations): Building and leading a global 24×7 cloud operations organization responsible for production monitoring, incident response, service restoration, and operational governance with an accent on automation, observability, and operational intelligence. Focus on reducing MTTD and MTTR, implementing AIOps and self-healing workflows, improving service availability, and establishing operational readiness across healthcare cloud platforms.

Location: Remote in the United States or hybrid, with hybrid preferred. Travel up to 20% domestically and internationally.

Company

hirify.global develops intelligent automation and cloud-native software to transform pharmacy and nursing care.

What you will do

  • Build and lead a global Site Reliability Operations organization supporting 24×7 production cloud services.
  • Own production operations, monitoring, event management, alert triage, operational escalation, and service health management.
  • Lead major incident management, executive and customer communications, service restoration, and post-incident improvement.
  • Establish operational readiness standards, including runbook validation, playbook development, monitoring validation, and disaster recovery readiness.
  • Drive automation and AI-assisted operations through intelligent alerting, self-healing workflows, automated runbooks, ChatOps, and predictive operations.
  • Govern operational processes, service levels, risk reviews, metrics, and strategic managed service provider performance.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field; a master’s degree is preferred.
  • 12+ years of experience in cloud operations, production operations, NOC, SRE, IT operations, or related disciplines.
  • 7+ years of progressive leadership experience managing managers and operational teams.
  • Experience leading 24×7 global production operations organizations.
  • Strong knowledge of cloud-native architectures, observability platforms, incident management, and operational governance.
  • Experience driving operational transformation through automation and process improvement, with strong executive communication and crisis leadership skills.

Nice to have

  • Experience building modern Site Reliability Operations or cloud operations organizations.
  • Experience with AIOps, observability platforms, and operational analytics.
  • Expertise with ITIL, Incident Management, Problem Management, and Change Management frameworks.
  • Experience supporting regulated healthcare, SaaS, or enterprise cloud platforms.
  • Experience managing strategic managed service providers.

Culture & Benefits

  • Global, distributed operations model with follow-the-sun coverage.
  • Culture centered on ownership, accountability, operational discipline, continuous improvement, and customer focus.
  • Close partnership with Site Reliability Engineering, Cloud Platform Engineering, Cloud Security, Product Engineering, Technical Support, and executive leadership.
  • Opportunity to develop operational leaders and shape a modern AI-enabled cloud operating model.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →