Назад
Company hidden
6 часов назад

Quality Engineering Manager (AI/SRE)

116 400 - 194 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Quality Engineering Manager (AI/SRE): Leading production support and reliability engineering for critical enterprise applications and cloud platforms with an accent on incident management, observability, automation, and Azure operations. Focus on resolving complex distributed-system issues, building self-healing support workflows, improving SLOs and resilience, and applying AI to troubleshooting and operational analytics.

Location: Buffalo, New York, United States of America

Salary: $116,400–$194,000 annual USD

Company

M&T Bank operates critical financial applications and platforms requiring secure, reliable, and compliant production operations.

What you will do

  • Serve as a technical escalation point for production incidents, outages, and service degradation.
  • Lead troubleshooting, root cause analysis, incident response, and post-incident improvements across applications, infrastructure, integrations, APIs, and cloud services.
  • Apply SRE practices, define SLIs and SLOs, improve resilience, and support disaster recovery and business continuity testing.
  • Design and optimize monitoring, logging, alerting, dashboards, distributed tracing, and other observability capabilities.
  • Develop automation, scripts, infrastructure-as-code workflows, and self-healing operational solutions.
  • Use AI-enabled tools for incident analysis, predictive analytics, knowledge management, and operational support.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • At least 5 years of experience supporting enterprise applications, platforms, or infrastructure in production environments.
  • Experience with complex troubleshooting, incident management, monitoring, observability, logging, alerting, scripting, and automation.
  • Strong analytical, problem-solving, communication, and collaboration skills.
  • Experience supporting cloud-hosted or hybrid environments, including Azure-based platforms.
  • Ability to work in Buffalo, New York, United States of America.

Nice to have

  • Experience with SRE practices, distributed systems, APIs, microservices, SLIs, SLOs, and production readiness reviews.
  • Experience with Dynatrace, Splunk, Datadog, Azure Monitor, Grafana, Prometheus, OpenTelemetry, or AppDynamics.
  • Advanced PowerShell, Python, REST API integration, infrastructure as code, CI/CD, and deployment pipelines.
  • Experience with Microsoft Copilot, Azure AI services, Microsoft Foundry, or AI-enabled support workflows.

Culture & Benefits

  • Cross-functional collaboration with Engineering, Product, Architecture, Infrastructure, Security, QA, and Vendor teams.
  • Emphasis on operational excellence, reliability engineering, continuous improvement, and customer experience.
  • Work aligned with enterprise risk, security, regulatory, audit, and governance requirements.
  • Opportunity to mentor support engineers and provide technical leadership during incident response.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →