обновлено 7 часов назад
Principal Site Reliability Engineer
139 700 - 232 900$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Principal Site Reliability Engineer (SRE/Cloud): Designing and improving highly reliable, scalable, and resilient enterprise platforms with an accent on observability, automation, incident management, and performance. Focus on defining SLOs and error budgets, building fault-tolerant architectures, leading root cause analysis and production readiness, and driving reliability standards across engineering teams.
Location: Buffalo, New York, United States of America
Salary: $139,700.00–$232,900.00 annual, USD
Company
M&T Bank operates in the financial services sector.
What you will do
- Define and drive service reliability standards, including SLOs, SLAs, and error budgets.
- Design highly available and fault-tolerant architectures that meet enterprise scalability and resiliency requirements.
- Lead incident management, problem management, root cause analysis, and post-incident reviews.
- Develop observability strategies covering logging, monitoring, alerting, and tracing.
- Lead automation initiatives for self-healing systems and operational workflows.
- Partner with development, infrastructure, cybersecurity, architecture, and senior stakeholder teams to improve reliability, performance, and capacity planning.
Requirements
- Associate’s degree and 9+ years of systems analysis and/or application development experience, or a bachelor’s degree and 7+ years of such experience.
- In lieu of a degree, 11+ years of combined education and/or relevant experience, including at least 7 years of systems analysis and/or application development experience.
- Expert experience in system design, reliability engineering, and production operations.
- Advanced proficiency in at least one programming or scripting language.
- Ability to apply SRE practices across multiple platforms and influence technical direction without direct authority.
Nice to have
- Experience with observability and incident management tooling.
- Experience with AWS or Azure.
- Strong understanding of CI/CD, DevOps, and SDLC practices.
- Experience defining and implementing SLO/SLI frameworks.
- Experience in regulated environments such as financial services.
Culture & Benefits
- Enterprise-wide reliability improvement initiatives across multiple platforms.
- Technical mentorship and leadership opportunities for less experienced engineers.
- Collaboration with development, infrastructure, cybersecurity, architecture, and senior stakeholders.
- Work aligned with risk, regulatory, internal control, and compliance standards.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Engineer (SRE/Incident Management)
100 000 - 215 000$
5 дней назад
Global Manager of Site Reliability Engineering (Java/AWS)
235 000 - 250 000$
4 дня назад
Senior Site Reliability Engineer (Fintech)
160 000 - 200 000$
3 дня назад
Senior Site Reliability Engineer (AI/Kubernetes)
137 900 - 221 400$
6 дней назад
Staff Engineer (SRE)
110 000 - 230 000$
5 дней назад