Назад
Company hidden
обновлено 7 часов назад

Principal Site Reliability Engineer

139 700 - 232 900$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Principal Site Reliability Engineer (SRE/Cloud): Designing and improving highly reliable, scalable, and resilient enterprise platforms with an accent on observability, automation, incident management, and performance. Focus on defining SLOs and error budgets, building fault-tolerant architectures, leading root cause analysis and production readiness, and driving reliability standards across engineering teams.

Location: Buffalo, New York, United States of America

Salary: $139,700.00–$232,900.00 annual, USD

Company

M&T Bank operates in the financial services sector.

What you will do

  • Define and drive service reliability standards, including SLOs, SLAs, and error budgets.
  • Design highly available and fault-tolerant architectures that meet enterprise scalability and resiliency requirements.
  • Lead incident management, problem management, root cause analysis, and post-incident reviews.
  • Develop observability strategies covering logging, monitoring, alerting, and tracing.
  • Lead automation initiatives for self-healing systems and operational workflows.
  • Partner with development, infrastructure, cybersecurity, architecture, and senior stakeholder teams to improve reliability, performance, and capacity planning.

Requirements

  • Associate’s degree and 9+ years of systems analysis and/or application development experience, or a bachelor’s degree and 7+ years of such experience.
  • In lieu of a degree, 11+ years of combined education and/or relevant experience, including at least 7 years of systems analysis and/or application development experience.
  • Expert experience in system design, reliability engineering, and production operations.
  • Advanced proficiency in at least one programming or scripting language.
  • Ability to apply SRE practices across multiple platforms and influence technical direction without direct authority.

Nice to have

  • Experience with observability and incident management tooling.
  • Experience with AWS or Azure.
  • Strong understanding of CI/CD, DevOps, and SDLC practices.
  • Experience defining and implementing SLO/SLI frameworks.
  • Experience in regulated environments such as financial services.

Culture & Benefits

  • Enterprise-wide reliability improvement initiatives across multiple platforms.
  • Technical mentorship and leadership opportunities for less experienced engineers.
  • Collaboration with development, infrastructure, cybersecurity, architecture, and senior stakeholders.
  • Work aligned with risk, regulatory, internal control, and compliance standards.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →