4 дня назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineering Manager (AWS/Kubernetes): Leading SRE operations for mission-critical deployments across on-premises DoD and AWS environments with an accent on team leadership, reliability planning, incident response, and operational ownership. Focus on coordinating infrastructure and application reliability work, guiding observability and automation improvements, and solving deployment and connectivity challenges in constrained and air-gapped environments.
Location: United States; remote-first hybrid role with regular on-site work at customer locations in Arlington, VA or Colorado Springs, CO. Candidates outside commuting distance must be willing to relocate.
Salary: $205,000–$255,000 per year, plus equity.
Company
builds collaboration and AI-powered workflow software for military planning and operational coordination.
What you will do
- Lead, hire, coach, and develop the SRE team across infrastructure, operations, and application reliability.
- Own capacity planning, prioritization, delivery commitments, and coordination across platform engineering, application engineering, security, and customer success.
- Set the reliability roadmap using customer needs, production data, incident patterns, and operational risks.
- Establish operational ownership, support responsibilities, escalation paths, and readiness criteria for deployments and releases.
- Support team-led incident response, sustainable on-call coverage, blameless postmortems, and follow-through on corrective actions.
- Guide infrastructure, automation, observability, security, and application reliability priorities across AWS, AWS GovCloud, on-premises, and air-gapped environments.
Requirements
- Active Secret clearance required.
- At least 5 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or a related infrastructure and operations role.
- Direct engineering management experience, including coaching, performance management, career development, and hiring.
- Experience with capacity planning, competing priorities, cross-team coordination, incident response, and post-incident reviews.
- Practical experience with Infrastructure as Code, configuration management, scripting, Kubernetes, CI/CD, AWS or AWS GovCloud, on-premises infrastructure, observability, networking, security, and application reliability.
- Must be able to work regularly on-site at customer locations; relocation assistance is available for candidates outside commuting distance.
Nice to have
- Experience supporting mission-critical customer deployments or coordinating delivery across multiple customer environments.
- Experience in DoD, classified, or air-gapped environments; familiarity with RMF, STIGs, or ICD 503.
- Experience with SLIs, SLOs, error budgets, GitOps, or on-premises virtualization platforms.
- Application development experience with TypeScript or Node.js.
- AWS DevOps Engineer, CKA, CKAD, Security+, or another DoD 8570.01-approved credential.
Culture & Benefits
- Remote-first organization with flexible work hours and unlimited PTO, while some customer-facing roles require on-site work.
- Health, dental, vision, and life insurance.
- 401(k) plan with company match and eight weeks of fully paid parental leave.
- Annual company summit trips and a $1,000 annual home office budget.
- Equity participation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Replit
5 дней назад
Staff Site Reliability Engineer (Kubernetes/GCP)
250 000 - 325 000$
Replit
5 дней назад
Site Reliability Engineer
210 000 - 275 000$
10 дней назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
5 дней назад
Site Reliability Engineer
105 600 - 145 200$
6 дней назад
Senior Site Reliability Engineer (ML)
200 000 - 225 000$
7 дней назад
Sr. Site Reliability Engineer
160 000 - 180 000$