8 часов назад
Site Reliability Engineering (SRE) Manager (Azure)
139 700 - 232 900$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineering (SRE) Manager (Azure): Leading SRE, production support, observability, incident management, and cloud reliability teams responsible for resilient business applications and platforms with an accent on operational excellence, automation, and production systems. Focus on defining SLOs, improving incident response and recovery, driving Azure modernization, and integrating AI-assisted operations while developing high-performing engineering teams.
Location: Buffalo, New York, United States of America
Salary: $139,700–$232,900 annually
Company
M&T Bank is a financial services organization operating critical business applications, platforms, and cloud services.
What you will do
- Lead SRE, production support, observability engineering, incident management, operational automation, cloud reliability, and platform operations teams.
- Define SRE strategies, SLIs, SLOs, operational health metrics, production readiness reviews, disaster recovery testing, and resilience assessments.
- Lead major incident response, executive communications, root cause analysis, service restoration, and corrective action tracking.
- Establish monitoring, alerting, logging, tracing, dashboards, and observability standards while reducing manual operational effort through automation.
- Partner with engineering and infrastructure teams on Azure, containers, APIs, microservices, hybrid environments, Infrastructure as Code, CI/CD, and platform engineering.
- Drive responsible AI and generative AI adoption for anomaly detection, automated diagnostics, incident response, and operational knowledge management.
Requirements
- 10+ years of technology experience in application support, infrastructure, cloud, software engineering, or reliability engineering.
- 5+ years of leadership experience managing engineering, operations, or SRE teams.
- Experience managing production systems supporting critical business functions and leading major incident response, root cause analysis, and service restoration.
- Strong knowledge of SRE principles, including SLOs, observability, automation, incident management, and operational excellence.
- Experience with cloud platforms, distributed systems, APIs, modern application architectures, and stakeholder management.
- Typically manages 10–20 direct and indirect reports.
Nice to have
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field.
- Experience with Azure cloud technologies, cloud-native architectures, and observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
- Experience with PowerShell, Python, Bash, APIs, CI/CD, Infrastructure as Code, DevOps, and platform engineering.
- Experience implementing operational AI use cases and working in financial services or another highly regulated industry.
Culture & Benefits
- Focus on ownership, accountability, innovation, collaboration, and continuous learning.
- Emphasis on reliability, customer experience, risk management, and continuous improvement.
- Market-informed annual compensation range of $139,700–$232,900.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Senior Site Reliability Engineer (AI Agents & Automation)
147 600 - 221 400$
4 дня назад
Software Engineering Manager-Site Reliability (SRE)
100 100 - 185 900$
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
5 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
4 дня назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
12 часов назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$