5 часов назад
Software Engineering Manager (Site Reliability Engineering)
110 110 - 204 490$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineering Manager (Site Reliability Engineering): Leading teams responsible for the reliability, scalability, and operational excellence of mission-critical platforms with an accent on incident management, observability, resiliency, and automation. Focus on resolving high-severity production incidents, improving distributed-system performance, overseeing disaster recovery and change execution, and reducing operational toil across a global 24x7 operation.
Location: Hybrid/in-office role based at a Technology Hub in Pittsburgh, Cleveland, Birmingham, Dallas, Phoenix, Denver, or another listed U.S. hub; weekly office attendance is required.
Base salary: $110,110–$204,490 per year, plus incentive eligibility.
Company
is a financial services corporation building and operating technology for customer-facing digital experiences.
What you will do
- Lead, coach, and develop SRE teams supporting mission-critical platforms and global 24x7 operations.
- Direct major incident response, real-time triage, remediation, stakeholder communication, and post-incident analysis.
- Drive root-cause analysis, permanent fixes, runbooks, knowledge sharing, and operational maturity improvements.
- Oversee production changes, release readiness, rollback planning, CAB reviews, and post-implementation reviews.
- Advance monitoring, alerting, dashboards, and observability using tools including Dynatrace, BigPanda, and LogScale.
- Improve availability, disaster recovery, scalability, performance, automation, governance, risk controls, and regulatory compliance.
Requirements
- 5+ years of related industry experience and 3+ years of management experience; a bachelor's degree or comparable combination of education and experience.
- Strong experience in Site Reliability Engineering, production support, or DevOps in high-availability enterprise environments.
- Deep knowledge of incident, problem, and change management frameworks, including major-incident response and root-cause analysis.
- Hands-on experience with monitoring tools, cloud or infrastructure platforms, automation, and reliability and observability improvements.
- Experience with Linux/Windows infrastructure, OCP, Oracle, SQL, MongoDB, and Cassandra; knowledge of Elasticsearch, Redis, MQ, and Kafka is beneficial.
- Availability for after-hours leadership and on-call support, including evenings, weekends, and holidays, may be required. does not provide employment visa sponsorship or participate in STEM OPT for this position.
Culture & Benefits
- Inclusive workplace focused on customer experience, operational excellence, ownership, and continuous learning.
- Full-time benefits may include medical, dental, vision, HSA, life and disability insurance, 401(k) matching, pension, and stock purchase plans.
- Educational assistance, wellness programs, dependent care support, and adoption, surrogacy, and doula reimbursement may be available.
- Paid time off may include parental leave, up to 11 holidays, occasional absence days, and 15–25 vacation days depending on career level and service.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Okta
8 часов назад
Manager, Site Reliability Engineering (Auth0) (Cloud Infrastructure)
182 000 - 250 800$
6 дней назад
Digital Shared Services Senior Manager (Testing and SRE)
167 600 - 279 400$
20 часов назад
Senior Manager, Site Reliability Engineering (SRE) (Fintech)
Affirm
6 дней назад
Manager, Software Engineering (Reliability Platform)
230 000 - 290 000$
3 дня назад
Sr. Director, Cloud and Infrastructure Transformation
214 900 - 358 100$
6 дней назад
Production Support Engineering Manager
140 000 - 170 000$