8 часов назад
Technical Engineer (Production Liability Engineer)
97 100 - 161 800$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Technical Engineer (Production Liability Engineer) (SRE/Cloud Operations): Ensuring the availability, stability, performance, and operational excellence of critical business applications and platforms with an accent on production support, observability, automation, Azure operations, and AI-assisted troubleshooting. Focus on resolving complex incidents, leading root cause analysis, improving system resilience, and eliminating operational toil through scripting and intelligent automation.
Location: Buffalo, New York, United States of America
Salary: $97,100–$161,800 annual (USD)
Company
M&T Bank is a banking organization supporting critical business applications, platforms, cloud services, and operational technology.
What you will do
- Serve as a technical escalation point for critical production incidents, outages, and service degradation.
- Lead troubleshooting, root cause analysis, incident response, and post-incident improvement activities across applications, infrastructure, integrations, APIs, and cloud services.
- Apply SRE practices, SLIs, SLOs, observability, monitoring, alerting, distributed tracing, and operational health metrics to improve reliability and resilience.
- Develop scripts, tools, workflows, infrastructure automation, diagnostics, health checks, remediation processes, and self-healing capabilities.
- Support Azure-hosted and hybrid environments, cloud deployments, platform upgrades, disaster recovery, failover, and business continuity testing.
- Build support runbooks and knowledge repositories while collaborating with Engineering, Architecture, Infrastructure, Security, Product, QA, and vendor teams.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- At least 5 years of experience supporting enterprise applications, platforms, or infrastructure in production environments.
- Experience troubleshooting complex application, infrastructure, integration, or cloud-related issues.
- Strong knowledge of incident management, production support, monitoring, observability, logging, alerting, scripting, and automation.
- Strong analytical, troubleshooting, communication, and collaboration skills.
- Must work in Buffalo, New York, United States of America.
Nice to have
- Experience with SRE practices, distributed systems, APIs, microservices, cloud-native applications, production readiness reviews, and operational metrics.
- Experience with Dynatrace, Splunk, Datadog, Azure Monitor, Grafana, Prometheus, OpenTelemetry, or AppDynamics.
- Advanced PowerShell, Python development, REST API integration, infrastructure as code, CI/CD tools, and deployment pipelines.
- Experience with Azure App Services, AKS, Functions, Storage, Event Hub, Service Bus, and cloud monitoring services.
- Experience using Microsoft Copilot, Azure AI services, Microsoft Foundry, or similar AI-enabled operational platforms.
Culture & Benefits
- Focus on operational excellence, system stability, resilience, performance, and customer and employee experience.
- Work across Engineering, Product, Architecture, Infrastructure, Security, QA, and vendor functions.
- Support enterprise risk, security, regulatory, audit, and operational governance requirements.
- Use automation and AI-enabled operational insights to reduce repetitive support work and improve troubleshooting.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Software Engineering Manager-Site Reliability (SRE)
100 100 - 185 900$
4 дня назад
Senior Site Reliability Engineer (AI Agents & Automation)
147 600 - 221 400$
1 день назад
Production Engineer (Fintech)
75 000 - 90 000$
5 дней назад
Principal Site Reliability Engineer (AI)
165 000 - 185 000$
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
7 дней назад