Назад
Company hidden
4 дня назад

Software Engineering Manager-Site Reliability (SRE)

100 100 - 185 900$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Software Engineering Manager-Site Reliability (SRE) (SRE, DevOps, Observability): Leading teams that maintain reliable, scalable, and operationally excellent platforms for critical digital services with an accent on incident management, resilience, monitoring, and automation. Focus on resolving high-severity production incidents, improving observability and availability, coordinating change and disaster recovery activities, and reducing operational toil across distributed systems.

Location: Hybrid role with weekly office attendance required at Technology Hub locations in Pittsburgh, Pennsylvania; Strongsville, Ohio; Birmingham, Alabama; Dallas, Texas; or Phoenix, Arizona.

Base salary: $100,100.00–$185,900.00 per year, plus incentive eligibility.

Company

hirify.global is a financial services company building and operating enterprise technology platforms for customer-facing digital experiences.

What you will do

  • Lead, coach, and develop SRE teams while aligning technology objectives with business needs.
  • Direct major incident response, real-time triage, remediation, post-incident analysis, and root-cause resolution.
  • Oversee production support across applications, Linux and Windows infrastructure, databases, middleware, integrations, and batch or ETL processes.
  • Improve monitoring, alerting, dashboards, and observability using platforms such as Dynatrace, BigPanda, and Logscale.
  • Drive high availability, disaster recovery, failover testing, scalability, performance optimization, and operational resilience.
  • Lead automation, change and release execution, governance, risk management, and 24x7 operational escalation.

Requirements

  • 5+ years of related industry experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, production support, or DevOps in high-availability enterprise environments.
  • Deep knowledge of incident, problem, and change management, including major-incident response and root-cause analysis.
  • Hands-on knowledge of monitoring tools, cloud or infrastructure platforms, automation, reliability engineering, and observability.
  • Experience with Linux or Windows, OCP, Oracle, SQL, MongoDB, and Cassandra; knowledge of Elasticsearch, Redis, MQ, and Kafka is beneficial.
  • Availability for after-hours, weekend, holiday, and on-call leadership support is required. A bachelor's degree or comparable education and experience is typically expected.

Nice to have

  • Working knowledge of Elasticsearch, Redis, MQ, and Kafka.
  • Experience with AI platforms, application development, release management, and IT automation.

Culture & Benefits

  • Inclusive, customer-focused workplace with emphasis on collaboration, ownership, continuous learning, and a blameless culture.
  • Medical, dental, vision, life, disability, HSA, 401(k) matching, pension, and stock purchase options.
  • Paid holidays, parental leave, vacation, occasional absence days, wellness programs, and educational assistance.
  • hirify.global does not provide employment visa sponsorship or participate in STEM OPT for this position.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →