Назад
9 часов назад

Manager, Site Reliability Engineering (Auth0) (Cloud Infrastructure)

182 000 - 250 800$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Manager, Site Reliability Engineering (Auth0) (Cloud Infrastructure): Leading the SRE team responsible for the reliability and operational excellence of Auth0 with an accent on scalable cloud-native infrastructure, observability, resilience, and incident response. Focus on designing monitoring and automation, troubleshooting critical systems in 24/7 on-call rotations, and mentoring engineers while driving architectural and cross-functional initiatives.

Location: New York, New York or Washington, DC; U.S. Person status required to access federal environments or protected federal data

Annual base salary: $182,000–$250,800 USD, with eligibility for equity, bonus, and benefits.

Company

Okta provides identity and access management infrastructure, including the Auth0 authentication platform.

What you will do

  • Lead the SRE team's technical direction and translate organizational goals into actionable roadmaps.
  • Drive complex, cross-functional reliability initiatives across product and platform teams.
  • Participate hands-on in 24/7 on-call rotations, troubleshoot incidents, and remediate issues on critical systems.
  • Design and implement monitoring, alerting, infrastructure automation, and resilience improvements.
  • Establish observability, resilience, blameless incident response, and software engineering practices.
  • Mentor SRE engineers and represent reliability in architectural reviews and strategic planning.

Requirements

  • 3+ years of hands-on team leadership in SRE or software engineering and 8+ years of total industry experience.
  • Deep experience with AWS or Azure, Terraform, containers, Kubernetes, microservices, and databases.
  • Strong programming skills in Go or Python, including production-grade automation and infrastructure tools.
  • Experience applying SRE principles, systematic problem-solving, observability, and incident response practices.
  • Strong verbal and written communication skills, including communication during high-pressure incidents.
  • Ability to submit documentation establishing U.S. Person status upon hire is required.

Nice to have

  • Experience improving uptime and reducing incident response times at scale.
  • Contributions to open-source infrastructure or observability tooling.
  • Experience designing incident response programs and runbook automation.

Culture & Benefits

  • Blameless incident response and a culture of continuous learning.
  • Remote-first environment with globally distributed teams.
  • Health, dental, and vision insurance, 401(k), flexible spending account, and paid leave including PTO and parental leave.
  • Equity and bonus opportunities where applicable.
  • In-person onboarding focused on connecting new hires with the mission and team.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →