6 дней назад
Software Engineering Manager - Site Reliability (SRE)
100 100 - 185 900$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Software Engineering Manager - Site Reliability (SRE) (Fintech/SRE): Leading teams responsible for the reliability, scalability, and operational excellence of mission-critical platforms with an accent on incident management, observability, automation, and production stability. Focus on resolving high-severity incidents, improving distributed-system resiliency, driving root-cause remediation, and coordinating 24x7 operational support.
Location: Hybrid role with weekly office time required in Pittsburgh, PA; Cleveland, OH; Birmingham, AL; Dallas, TX; or Phoenix, AZ.
Base salary: $100,100–$185,900 per year, plus incentive eligibility.
Company
is a financial services organization focused on delivering secure, reliable digital experiences and maintaining strong risk and operational controls.
What you will do
- Lead, coach, and develop SRE teams supporting mission-critical platforms and distributed systems.
- Provide leadership during major incidents, production escalations, after-hours events, and on-call rotations.
- Drive incident response, root-cause analysis, problem management, remediation, and continuous reliability improvements.
- Improve monitoring, alerting, dashboards, and observability using tools such as Dynatrace, BigPanda, and Logscale.
- Lead change and release execution, disaster recovery, failover testing, capacity planning, and performance optimization.
- Reduce operational toil through automation, standardized runbooks, and improved production-support processes.
Requirements
- 5+ years of related industry experience and 3+ years of management experience.
- Strong experience in Site Reliability Engineering, production support, or DevOps in high-availability enterprise environments.
- Experience with incident, problem, change, and release management frameworks.
- Hands-on knowledge of monitoring tools, cloud or infrastructure platforms, automation, and observability practices.
- Experience with Linux or Windows infrastructure, OCP, Oracle, SQL, MongoDB, and Cassandra; knowledge of Elasticsearch, Redis, MQ, or Kafka is beneficial.
- will not provide employment visa sponsorship or participate in STEM OPT for this position.
Nice to have
- Experience with Elasticsearch, Redis, MQ, and Kafka.
- Experience improving operational maturity, service performance, and automation across complex enterprise systems.
Culture & Benefits
- In-office company culture with a supportive and inclusive workplace.
- Full-time benefits may include medical, dental, vision, life and disability insurance, HSA options, 401(k) matching, pension, and stock purchase plans.
- Educational assistance, wellness programs, dependent-care support, and family-related reimbursements are available depending on eligibility.
- Paid time off includes holidays, occasional absence days, vacation, and parental or maternity leave depending on eligibility.
- Participation in leadership escalation and on-call rotations supports a global 24x7 operation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Manager of Site Reliability Engineering (Azure/AWS)
160 000 - 180 000$
8 дней назад
Senior Manager, Site Reliability Engineering (AI Ops)
222 000 - 300 500$
9 дней назад
Senior Site Reliability Engineering
182 800 - 247 300$
11 дней назад
Manager, Site Reliability Engineering (Azure)
139 700 - 232 900$
9 дней назад
Software Engineer (Cloud Infrastructure/SRE)
147 900 - 220 000$
9 дней назад
Staff Site Reliability Engineer (Autonomous Vehicles)
172 000 - 300 000$