Назад
Company hidden
8 часов назад

Site Reliability Engineering (SRE) Manager (Azure)

139 700 - 232 900$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineering (SRE) Manager (Azure): Leading SRE, production support, observability, incident management, and cloud reliability teams responsible for resilient business applications and platforms with an accent on operational excellence, automation, and production systems. Focus on defining SLOs, improving incident response and recovery, driving Azure modernization, and integrating AI-assisted operations while developing high-performing engineering teams.

Location: Buffalo, New York, United States of America

Salary: $139,700–$232,900 annually

Company

M&T Bank is a financial services organization operating critical business applications, platforms, and cloud services.

What you will do

  • Lead SRE, production support, observability engineering, incident management, operational automation, cloud reliability, and platform operations teams.
  • Define SRE strategies, SLIs, SLOs, operational health metrics, production readiness reviews, disaster recovery testing, and resilience assessments.
  • Lead major incident response, executive communications, root cause analysis, service restoration, and corrective action tracking.
  • Establish monitoring, alerting, logging, tracing, dashboards, and observability standards while reducing manual operational effort through automation.
  • Partner with engineering and infrastructure teams on Azure, containers, APIs, microservices, hybrid environments, Infrastructure as Code, CI/CD, and platform engineering.
  • Drive responsible AI and generative AI adoption for anomaly detection, automated diagnostics, incident response, and operational knowledge management.

Requirements

  • 10+ years of technology experience in application support, infrastructure, cloud, software engineering, or reliability engineering.
  • 5+ years of leadership experience managing engineering, operations, or SRE teams.
  • Experience managing production systems supporting critical business functions and leading major incident response, root cause analysis, and service restoration.
  • Strong knowledge of SRE principles, including SLOs, observability, automation, incident management, and operational excellence.
  • Experience with cloud platforms, distributed systems, APIs, modern application architectures, and stakeholder management.
  • Typically manages 10–20 direct and indirect reports.

Nice to have

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related field.
  • Experience with Azure cloud technologies, cloud-native architectures, and observability platforms such as Dynatrace, Splunk, Datadog, Grafana, Azure Monitor, or OpenTelemetry.
  • Experience with PowerShell, Python, Bash, APIs, CI/CD, Infrastructure as Code, DevOps, and platform engineering.
  • Experience implementing operational AI use cases and working in financial services or another highly regulated industry.

Culture & Benefits

  • Focus on ownership, accountability, innovation, collaboration, and continuous learning.
  • Emphasis on reliability, customer experience, risk management, and continuous improvement.
  • Market-informed annual compensation range of $139,700–$232,900.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →