4 дня назад
Software Engineering Manager-Site Reliability (SRE)
100 100 - 185 900$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineering Manager-Site Reliability (SRE) (SRE, DevOps, Observability): Leading teams that maintain reliable, scalable, and operationally excellent platforms for critical digital services with an accent on incident management, resilience, monitoring, and automation. Focus on resolving high-severity production incidents, improving observability and availability, coordinating change and disaster recovery activities, and reducing operational toil across distributed systems.
Location: Hybrid role with weekly office attendance required at Technology Hub locations in Pittsburgh, Pennsylvania; Strongsville, Ohio; Birmingham, Alabama; Dallas, Texas; or Phoenix, Arizona.
Base salary: $100,100.00–$185,900.00 per year, plus incentive eligibility.
Company
is a financial services company building and operating enterprise technology platforms for customer-facing digital experiences.
What you will do
- Lead, coach, and develop SRE teams while aligning technology objectives with business needs.
- Direct major incident response, real-time triage, remediation, post-incident analysis, and root-cause resolution.
- Oversee production support across applications, Linux and Windows infrastructure, databases, middleware, integrations, and batch or ETL processes.
- Improve monitoring, alerting, dashboards, and observability using platforms such as Dynatrace, BigPanda, and Logscale.
- Drive high availability, disaster recovery, failover testing, scalability, performance optimization, and operational resilience.
- Lead automation, change and release execution, governance, risk management, and 24x7 operational escalation.
Requirements
- 5+ years of related industry experience and 3+ years of management experience.
- Strong experience in Site Reliability Engineering, production support, or DevOps in high-availability enterprise environments.
- Deep knowledge of incident, problem, and change management, including major-incident response and root-cause analysis.
- Hands-on knowledge of monitoring tools, cloud or infrastructure platforms, automation, reliability engineering, and observability.
- Experience with Linux or Windows, OCP, Oracle, SQL, MongoDB, and Cassandra; knowledge of Elasticsearch, Redis, MQ, and Kafka is beneficial.
- Availability for after-hours, weekend, holiday, and on-call leadership support is required. A bachelor's degree or comparable education and experience is typically expected.
Nice to have
- Working knowledge of Elasticsearch, Redis, MQ, and Kafka.
- Experience with AI platforms, application development, release management, and IT automation.
Culture & Benefits
- Inclusive, customer-focused workplace with emphasis on collaboration, ownership, continuous learning, and a blameless culture.
- Medical, dental, vision, life, disability, HSA, 401(k) matching, pension, and stock purchase options.
- Paid holidays, parental leave, vacation, occasional absence days, wellness programs, and educational assistance.
- does not provide employment visa sponsorship or participate in STEM OPT for this position.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Principal Site Reliability Engineer (Paze)
194 000 - 237 000$
6 дней назад
Site Reliability Engineer (DevOps)
70 000 - 75 000$
Replit
5 дней назад
Site Reliability Engineer
210 000 - 275 000$
8 дней назад
Staff Site Reliability Engineer (Paze)
157 000 - 192 000$
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$
11 дней назад
Senior Staff Site Reliability Engineer
232 338 - 290 422$