2 дня назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes): Owning reliability, scalability, and security for production applications and platforms across on-premise DoD environments and AWS cloud with an accent on observability, incident response, and infrastructure automation. Focus on designing Kubernetes and Infrastructure-as-Code solutions, defining SLIs and SLOs, leading blameless postmortems, and building resilient systems for air-gapped and sensitive environments.
Location: United States; remote work with regular on-site work at customer locations in Arlington, VA. Candidates outside commuting distance must be willing to relocate to the United States.
Salary: $180K–$220K per year, plus equity.
Company
develops collaboration and AI-powered workflow software for military planning and operational coordination.
What you will do
- Own the reliability, scalability, and security of production applications and platforms across on-premise DoD environments and AWS cloud.
- Design and manage monitoring, logging, alerting, and observability using tools such as Prometheus, Loki, Alloy, and Grafana.
- Define and maintain service level indicators and objectives, including actionable alerting and error budgets.
- Lead incident response, critical incident coordination, root-cause analysis, and blameless post-incident reviews.
- Build secure Kubernetes clusters and resilient cloud and on-premise environments with Terraform, Ansible, and embedded compliance controls.
- Automate operational work, reduce toil, and improve deployment and production readiness for air-gapped environments.
Requirements
- Active Top Secret clearance required; SCI eligibility is a plus.
- Regular on-site work at customer locations in Arlington, VA is required; relocation assistance is available for candidates who are not within commuting distance.
- 5+ years of experience in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
- Experience with Terraform or CloudFormation, Ansible, Kubernetes, and CI/CD pipelines such as GitLab CI/CD, Jenkins, or GitHub Actions.
- Proficiency with at least one of Python, Go, or Bash, plus familiarity with AWS or AWS GovCloud and secure networking fundamentals.
- Experience with incident response, root-cause analysis, and collaboration across platform, DevOps, and application teams.
Nice to have
- DoD environments and compliance frameworks such as RMF, STIGs, or ICD 503.
- GitOps, security-minded design, and meaningful SLIs/SLOs for distributed systems.
- On-premise virtualization with VMware, Proxmox, Nutanix, or Hyper-V.
- Service mesh experience with Istio or Linkerd and relevant AWS, Kubernetes, or DoD security certifications.
Culture & Benefits
- Remote-first organization with flexible work hours and unlimited PTO, while some customer-facing roles require on-site work.
- Health, dental, vision, and life insurance.
- 401(k) plan with company match and eight weeks of fully paid parental leave.
- Annual company retreats and a $1,000 yearly home office budget.
- Equity participation and a culture of blameless postmortems, mentoring, and continuous reliability improvement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$
5 дней назад
Site Reliability Engineer (AWS)
120 000 - 185 000$
7 дней назад
Site Reliability Engineer (Azure/Terraform)
105 600 - 145 200$
2 дня назад
Senior Site Reliability Engineer (Cloud Infrastructure)
145 000 - 175 000$
4 дня назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
6 дней назад
Senior Site Reliability Engineer (Azure/AWS)
130 000 - 160 000$