7 часов назад
Manager, Site Reliability Engineering (Azure)
139 700 - 232 900$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Manager, Site Reliability Engineering (Azure): Leading an enterprise SRE Center of Excellence and Forward Deployed SRE program supporting critical banking platforms, applications, and technology services with an accent on reliability strategy, observability, release engineering, and operational resilience. Focus on establishing SLO and error-budget governance, improving progressive deployments and automated recovery, and coordinating reliability outcomes across engineering, cloud, operations, risk, and business teams.
Location: Buffalo, New York, United States of America
Salary: $139,700.00–$232,900.00 annual USD
Company
M&T Bank is a banking organization seeking enterprise technology leadership for reliability, resilience, and operational maturity across critical banking platforms and services.
What you will do
- Lead the Site Reliability Engineering Center of Excellence and Forward Deployed SRE program across application, platform, cloud, infrastructure, and operations teams.
- Define the enterprise SRE strategy, operating model, standards, governance, service offerings, and multiyear adoption roadmap.
- Manage people leaders, program managers, technical leads, senior engineers, employees, and contingent resources.
- Establish governance for SLOs, SLIs, error budgets, operational readiness, reliability metrics, and production risk.
- Improve release and deployment safety through progressive delivery, CI/CD integration, automated validation, rollback, and deployment observability.
- Drive observability, automation, Infrastructure as Code, incident response, resiliency testing, toil reduction, and measurable service improvements.
Requirements
- At least 11 years of combined higher education and/or work experience, including 4 years of engineering or architecture experience and 5 years of leadership experience with people management.
- Enterprise leadership experience in SRE, reliability engineering, production engineering, or comparable technology capabilities.
- Strong knowledge of SRE practices, SLOs, SLIs, error budgets, observability, incident and problem management, resiliency, capacity planning, and toil reduction.
- Experience with cloud-native and hybrid architectures, distributed systems, APIs, containers, Microsoft Azure, CI/CD, and Infrastructure as Code, preferably Terraform.
- Experience with progressive delivery, blue-green and canary deployments, automated health validation, rollback, release quality gates, and disaster recovery.
- Ability to lead multidisciplinary teams across technology domains and present strategy, performance, investment needs, and risks to senior management.
Nice to have
- Bachelor’s or Master’s degree in engineering, computer science, information systems, business administration, or a related discipline.
- At least 12 years of technology management or large program leadership experience.
- Experience leading an SRE Center of Excellence, production engineering organization, platform engineering function, or Forward Deployed SRE model.
- Exposure to AI-assisted operations, AIOps, automated incident correlation, predictive analytics, performance engineering, or chaos engineering.
Culture & Benefits
- Full-time employment with compensation aligned to market-informed pay practices.
- Work spans multiple technology organizations, geographies, and time zones.
- Emphasis on engineering discipline, continuous learning, accountability, collaboration, knowledge sharing, and belonging.
- Responsibility for regulatory compliance, cybersecurity, technology risk, audit support, and control effectiveness in a banking environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →