6 часов назад
Quality Engineering Manager (AI/SRE)
116 400 - 194 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Quality Engineering Manager (AI/SRE): Leading production support and reliability engineering for critical enterprise applications and cloud platforms with an accent on incident management, observability, automation, and Azure operations. Focus on resolving complex distributed-system issues, building self-healing support workflows, improving SLOs and resilience, and applying AI to troubleshooting and operational analytics.
Location: Buffalo, New York, United States of America
Salary: $116,400–$194,000 annual USD
Company
M&T Bank operates critical financial applications and platforms requiring secure, reliable, and compliant production operations.
What you will do
- Serve as a technical escalation point for production incidents, outages, and service degradation.
- Lead troubleshooting, root cause analysis, incident response, and post-incident improvements across applications, infrastructure, integrations, APIs, and cloud services.
- Apply SRE practices, define SLIs and SLOs, improve resilience, and support disaster recovery and business continuity testing.
- Design and optimize monitoring, logging, alerting, dashboards, distributed tracing, and other observability capabilities.
- Develop automation, scripts, infrastructure-as-code workflows, and self-healing operational solutions.
- Use AI-enabled tools for incident analysis, predictive analytics, knowledge management, and operational support.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- At least 5 years of experience supporting enterprise applications, platforms, or infrastructure in production environments.
- Experience with complex troubleshooting, incident management, monitoring, observability, logging, alerting, scripting, and automation.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Experience supporting cloud-hosted or hybrid environments, including Azure-based platforms.
- Ability to work in Buffalo, New York, United States of America.
Nice to have
- Experience with SRE practices, distributed systems, APIs, microservices, SLIs, SLOs, and production readiness reviews.
- Experience with Dynatrace, Splunk, Datadog, Azure Monitor, Grafana, Prometheus, OpenTelemetry, or AppDynamics.
- Advanced PowerShell, Python, REST API integration, infrastructure as code, CI/CD, and deployment pipelines.
- Experience with Microsoft Copilot, Azure AI services, Microsoft Foundry, or AI-enabled support workflows.
Culture & Benefits
- Cross-functional collaboration with Engineering, Product, Architecture, Infrastructure, Security, QA, and Vendor teams.
- Emphasis on operational excellence, reliability engineering, continuous improvement, and customer experience.
- Work aligned with enterprise risk, security, regulatory, audit, and governance requirements.
- Opportunity to mentor support engineers and provide technical leadership during incident response.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Software Engineering Manager-Site Reliability (SRE)
100 100 - 185 900$
5 дней назад
Site Reliability Engineering Manager (Fintech)
250 000 - 280 000$
5 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
1 день назад
Production Support Engineering Manager
140 000 - 170 000$
4 дня назад
Engineering Manager, Infrastructure (AI)
156 000 - 224 250$
4 дня назад
Senior Site Reliability Engineer (AI)
185 500 - 232 000$