10 дней назад
Senior Manager - Site Reliability Engineering (SRE)
9 458 - 16 551$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Manager - Site Reliability Engineering (SRE) (Linux/Kubernetes): Leading SRE teams and modernizing the reliability, automation, observability, and performance of Linux-based digital commerce infrastructure with an accent on Kubernetes, CI/CD, Terraform, incident management, and team development. Focus on defining SRE strategy, building durable improvements from incident analysis, and balancing reliability, delivery speed, and technical debt across a large commerce platform.
Location: Remote within the Eastern or Central Time Zones, or hybrid in Newport News, Virginia
Salary: $9,458.97–$16,551.03 per month; bonus or incentive plan eligible
Company
is a Fortune 500 supplier serving commercial, residential, industrial, facilities, HVAC, waterworks, and digital commerce markets, with approximately 36,000 associates across 1,700 locations.
What you will do
- Lead and develop a high-performing Site Reliability Engineering team responsible for the reliability, availability, and performance of Linux-based digital commerce infrastructure.
- Set the technical and strategic direction for SRE, including Kubernetes, Docker, modern DevOps practices, and platform modernization.
- Establish CI/CD and infrastructure-as-code standards using GitHub Actions, Terraform, and related tooling.
- Define configuration management and automation standards with Puppet and Python, reducing operational toil.
- Guide strategies for load balancing, application delivery, virtualization, web and application server performance, observability, and artifact management.
- Own incident management, on-call practices, root cause analysis, SLOs, SLIs, error budgets, and cross-functional reliability initiatives.
Requirements
- Applicants must be based within the Eastern or Central Time Zones.
- 8+ years of professional Linux systems administration experience in production environments, including 3+ years leading or managing engineering or SRE teams.
- Strong experience with Kubernetes, Docker or similar container platforms, CI/CD pipelines, GitHub Actions, and Terraform or comparable infrastructure-as-code tools.
- Extensive experience with Puppet, Python or a similar language, load balancing, application delivery, virtualization, Nginx, and Apache Tomcat.
- Experience with Datadog, JFrog Artifactory, networking, storage, enterprise infrastructure security, incident response, on-call structures, SLIs, SLOs, and error budgets.
- Strong systems-thinking, troubleshooting, communication, leadership, mentoring, and cross-functional collaboration skills.
Culture & Benefits
- Remote or hybrid work options in accordance with company policy.
- Health, dental, and vision insurance, paid time off, life insurance, and a 401(k) with company match.
- Mental health coverage, gender-affirming and family-building benefits, and paid parental leave.
- Associate discounts and community involvement opportunities.
- Focus on accountability, collaboration, innovation, continuous learning, engineering excellence, and customer outcomes.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
13 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
13 дней назад
Site Reliability Engineer
90 000 - 110 000$
13 дней назад
Software Development Engineer, SRE (US Federal)
137 000 - 205 400$
12 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
2 дня назад
Senior Site Reliability Engineer (AI/Kubernetes)
137 900 - 221 400$
13 дней назад
Staff Site Reliability Engineer (Production Engineer) - Federal
119 000 - 170 000$