2 дня назад
Site Reliability Engineer (AWS)
180 000 - 220 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AWS): Operating and improving mission-critical deployments across on-premise DoD environments and AWS cloud with an accent on observability, incident response, and secure infrastructure automation. Focus on designing Kubernetes platforms, defining SLIs and SLOs, embedding RMF and STIG controls, and eliminating operational toil in air-gapped environments.
Location: Regular on-site work at customer locations in Colorado Springs, Colorado. Candidates outside commuting distance must be willing to relocate; relocation assistance is provided.
Salary: $180,000–$220,000 per year, plus equity.
Company
builds collaboration and AI-powered workflow software for military planning and operational coordination.
What you will do
- Operate and support mission-critical deployments across on-premise DoD environments and AWS cloud environments.
- Design and manage monitoring, logging, and alerting with observability tools such as Prometheus, Loki, Alloy, and Grafana.
- Define and measure SLIs and SLOs, improve reliability practices, and establish actionable alerting.
- Lead incident response, root-cause analysis, blameless postmortems, and After Action Reviews.
- Build secure, resilient Kubernetes clusters and cloud/on-premise environments using Terraform and Ansible.
- Automate operational work, embed RMF and STIG controls, and improve deployment and management of on-premise environments.
Requirements
- Active Top Secret clearance required; SCI eligibility is a plus.
- 5+ years of experience in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
- Experience with Terraform or CloudFormation, Ansible, Kubernetes, and CI/CD pipelines using GitLab CI/CD, Jenkins, or GitHub Actions.
- Proficiency in at least one of Python, Go, or Bash, plus familiarity with AWS or AWS GovCloud.
- Experience with observability platforms such as the Grafana or ELK stacks or Datadog.
- Strong incident response, root-cause analysis, networking, collaboration, and continuous-improvement skills.
Nice to have
- DoD environment experience and familiarity with RMF, STIGs, or ICD 503.
- GitOps practices, sensitive-environment security design, and meaningful SLI/SLO or error-budget implementation.
- On-premise virtualization with VMware, Proxmox, Nutanix, or Hyper-V.
- Service mesh experience with Istio or Linkerd and relevant AWS or Kubernetes certifications.
- Active Security+ or another DoD 8570.01-approved credential, or the ability to obtain one within three months.
Culture & Benefits
- Remote-first organization with flexible work hours, while this role requires regular customer-site presence.
- Unlimited PTO, health, dental, vision, and life insurance.
- 401(k) plan with company match and eight weeks of fully paid parental leave.
- Annual company summit trips and a $1,000 annual home-office budget.
- Equity participation and a culture of blameless postmortems, mentoring, and proactive reliability improvement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 часов назад
DevOps Engineer (AWS)
115 199 - 160 000$
6 дней назад
DevOps Engineer (AWS)
100 173 - 135 000$
5 дней назад
Senior DevOps / Cloud Infrastructure Engineer (AWS)
115 000 - 140 000$
12 часов назад
DevOps Engineer (AWS/Kubernetes)
100 173 - 130 000$
11 часов назад
Member of Technical Staff, Infrastructure Engineer (AI)
175 000 - 240 000$
2 дня назад
DevOps Engineer (AI)
163 000 - 204 000$