9 часов назад
Manager, Site Reliability Engineering (Auth0) (Cloud Infrastructure)
182 000 - 250 800$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Manager, Site Reliability Engineering (Auth0) (Cloud Infrastructure): Leading the SRE team responsible for the reliability and operational excellence of Auth0 with an accent on scalable cloud-native infrastructure, observability, resilience, and incident response. Focus on designing monitoring and automation, troubleshooting critical systems in 24/7 on-call rotations, and mentoring engineers while driving architectural and cross-functional initiatives.
Location: New York, New York or Washington, DC; U.S. Person status required to access federal environments or protected federal data
Annual base salary: $182,000–$250,800 USD, with eligibility for equity, bonus, and benefits.
Company
Okta provides identity and access management infrastructure, including the Auth0 authentication platform.
What you will do
- Lead the SRE team's technical direction and translate organizational goals into actionable roadmaps.
- Drive complex, cross-functional reliability initiatives across product and platform teams.
- Participate hands-on in 24/7 on-call rotations, troubleshoot incidents, and remediate issues on critical systems.
- Design and implement monitoring, alerting, infrastructure automation, and resilience improvements.
- Establish observability, resilience, blameless incident response, and software engineering practices.
- Mentor SRE engineers and represent reliability in architectural reviews and strategic planning.
Requirements
- 3+ years of hands-on team leadership in SRE or software engineering and 8+ years of total industry experience.
- Deep experience with AWS or Azure, Terraform, containers, Kubernetes, microservices, and databases.
- Strong programming skills in Go or Python, including production-grade automation and infrastructure tools.
- Experience applying SRE principles, systematic problem-solving, observability, and incident response practices.
- Strong verbal and written communication skills, including communication during high-pressure incidents.
- Ability to submit documentation establishing U.S. Person status upon hire is required.
Nice to have
- Experience improving uptime and reducing incident response times at scale.
- Contributions to open-source infrastructure or observability tooling.
- Experience designing incident response programs and runbook automation.
Culture & Benefits
- Blameless incident response and a culture of continuous learning.
- Remote-first environment with globally distributed teams.
- Health, dental, and vision insurance, 401(k), flexible spending account, and paid leave including PTO and parental leave.
- Equity and bonus opportunities where applicable.
- In-person onboarding focused on connecting new hires with the mission and team.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 часов назад
Engineering Manager, AgentControl (AI)
163 000 - 263 670$
12 часов назад
Director of Engineering, Cloud & Reliability (AWS)
209 500 - 270 000$
Affirm
6 дней назад
Manager, Software Engineering (Reliability Platform)
230 000 - 290 000$
Affirm
5 дней назад
Manager, Software Engineering (Infra Foundations)
181 000 - 241 000CAD
7 дней назад
Director Of Software Engineering and Architecture (AI)
232 100 - 264 200$
6 часов назад
Software Engineering Manager (Site Reliability Engineering)
110 110 - 204 490$