Назад
Company hidden
8 часов назад

Principal Site Reliability Engineer

139 700 - 232 900$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Site Reliability Engineer (SRE/DevOps): Designing and improving reliable, scalable, and resilient platform solutions across the enterprise with an accent on service reliability standards, observability, automation, and production operations. Focus on building fault-tolerant architectures, leading incident and problem management, implementing self-healing systems, and driving enterprise-wide reliability improvements in a regulated financial services environment.

Location: Buffalo, New York, United States of America

Salary: $139,700–$232,900 annual (USD)

Company

M&T Bank is a financial services organization operating in a regulated environment.

What you will do

  • Define and drive service reliability standards, including SLOs, SLAs, and error budgets.
  • Design highly available and fault-tolerant architectures aligned with scalability and resiliency requirements.
  • Lead incident detection, response, escalation, recovery, post-incident reviews, and root cause analysis.
  • Develop observability strategies covering logging, monitoring, alerting, and tracing.
  • Lead automation initiatives for self-healing systems and operational workflows.
  • Partner with development, infrastructure, cybersecurity, and architecture teams while driving cross-team reliability improvements and mentoring engineers.

Requirements

  • Associate’s degree with at least 9 years of systems analysis and/or application development experience, or a bachelor’s degree with at least 7 years of relevant experience.
  • In lieu of a degree, at least 11 years of combined education and/or relevant work experience, including at least 7 years of systems analysis and/or application development experience.
  • Expert experience in system design, reliability engineering, and production operations.
  • Advanced proficiency in at least one programming or scripting language.
  • Experience with observability and incident management tooling, SLO/SLI frameworks, and production readiness practices.
  • Strong understanding of CI/CD, DevOps, SDLC, performance, resilience, capacity planning, risk, and regulatory standards.

Nice to have

  • Experience with cloud platforms such as AWS or Azure.
  • Experience in regulated environments such as financial services.
  • Strong communication and stakeholder management skills.

Culture & Benefits

  • Enterprise-wide technical influence across multiple platforms.
  • Opportunity to lead complex reliability initiatives without direct supervisory responsibility.
  • Work includes mentoring and technical leadership across Technology.
  • Compensation is informed by the candidate’s knowledge, skills, and experience within the stated annual pay range.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →