Назад
Company hidden
10 дней назад

Senior Manager - Site Reliability Engineering (SRE)

9 458 - 16 551$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Manager - Site Reliability Engineering (SRE) (Linux/Kubernetes): Leading SRE teams and modernizing the reliability, automation, observability, and performance of Linux-based digital commerce infrastructure with an accent on Kubernetes, CI/CD, Terraform, incident management, and team development. Focus on defining SRE strategy, building durable improvements from incident analysis, and balancing reliability, delivery speed, and technical debt across a large commerce platform.

Location: Remote within the Eastern or Central Time Zones, or hybrid in Newport News, Virginia

Salary: $9,458.97–$16,551.03 per month; bonus or incentive plan eligible

Company

hirify.global is a Fortune 500 supplier serving commercial, residential, industrial, facilities, HVAC, waterworks, and digital commerce markets, with approximately 36,000 associates across 1,700 locations.

What you will do

  • Lead and develop a high-performing Site Reliability Engineering team responsible for the reliability, availability, and performance of Linux-based digital commerce infrastructure.
  • Set the technical and strategic direction for SRE, including Kubernetes, Docker, modern DevOps practices, and platform modernization.
  • Establish CI/CD and infrastructure-as-code standards using GitHub Actions, Terraform, and related tooling.
  • Define configuration management and automation standards with Puppet and Python, reducing operational toil.
  • Guide strategies for load balancing, application delivery, virtualization, web and application server performance, observability, and artifact management.
  • Own incident management, on-call practices, root cause analysis, SLOs, SLIs, error budgets, and cross-functional reliability initiatives.

Requirements

  • Applicants must be based within the Eastern or Central Time Zones.
  • 8+ years of professional Linux systems administration experience in production environments, including 3+ years leading or managing engineering or SRE teams.
  • Strong experience with Kubernetes, Docker or similar container platforms, CI/CD pipelines, GitHub Actions, and Terraform or comparable infrastructure-as-code tools.
  • Extensive experience with Puppet, Python or a similar language, load balancing, application delivery, virtualization, Nginx, and Apache Tomcat.
  • Experience with Datadog, JFrog Artifactory, networking, storage, enterprise infrastructure security, incident response, on-call structures, SLIs, SLOs, and error budgets.
  • Strong systems-thinking, troubleshooting, communication, leadership, mentoring, and cross-functional collaboration skills.

Culture & Benefits

  • Remote or hybrid work options in accordance with company policy.
  • Health, dental, and vision insurance, paid time off, life insurance, and a 401(k) with company match.
  • Mental health coverage, gender-affirming and family-building benefits, and paid parental leave.
  • Associate discounts and community involvement opportunities.
  • Focus on accountability, collaboration, innovation, continuous learning, engineering excellence, and customer outcomes.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →