Назад
Company hidden
2 дня назад

Site Reliability Engineer (Kubernetes)

180 000 - 220 000$
Формат работы
remote (только USA)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Релокация
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Kubernetes): Owning reliability, scalability, and security for production applications and platforms across on-premise DoD environments and AWS cloud with an accent on observability, incident response, and infrastructure automation. Focus on designing Kubernetes and Infrastructure-as-Code solutions, defining SLIs and SLOs, leading blameless postmortems, and building resilient systems for air-gapped and sensitive environments.

Location: United States; remote work with regular on-site work at customer locations in Arlington, VA. Candidates outside commuting distance must be willing to relocate to the United States.

Salary: $180K–$220K per year, plus equity.

Company

hirify.global develops collaboration and AI-powered workflow software for military planning and operational coordination.

What you will do

  • Own the reliability, scalability, and security of production applications and platforms across on-premise DoD environments and AWS cloud.
  • Design and manage monitoring, logging, alerting, and observability using tools such as Prometheus, Loki, Alloy, and Grafana.
  • Define and maintain service level indicators and objectives, including actionable alerting and error budgets.
  • Lead incident response, critical incident coordination, root-cause analysis, and blameless post-incident reviews.
  • Build secure Kubernetes clusters and resilient cloud and on-premise environments with Terraform, Ansible, and embedded compliance controls.
  • Automate operational work, reduce toil, and improve deployment and production readiness for air-gapped environments.

Requirements

  • Active Top Secret clearance required; SCI eligibility is a plus.
  • Regular on-site work at customer locations in Arlington, VA is required; relocation assistance is available for candidates who are not within commuting distance.
  • 5+ years of experience in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
  • Experience with Terraform or CloudFormation, Ansible, Kubernetes, and CI/CD pipelines such as GitLab CI/CD, Jenkins, or GitHub Actions.
  • Proficiency with at least one of Python, Go, or Bash, plus familiarity with AWS or AWS GovCloud and secure networking fundamentals.
  • Experience with incident response, root-cause analysis, and collaboration across platform, DevOps, and application teams.

Nice to have

  • DoD environments and compliance frameworks such as RMF, STIGs, or ICD 503.
  • GitOps, security-minded design, and meaningful SLIs/SLOs for distributed systems.
  • On-premise virtualization with VMware, Proxmox, Nutanix, or Hyper-V.
  • Service mesh experience with Istio or Linkerd and relevant AWS, Kubernetes, or DoD security certifications.

Culture & Benefits

  • Remote-first organization with flexible work hours and unlimited PTO, while some customer-facing roles require on-site work.
  • Health, dental, vision, and life insurance.
  • 401(k) plan with company match and eight weeks of fully paid parental leave.
  • Annual company retreats and a $1,000 yearly home office budget.
  • Equity participation and a culture of blameless postmortems, mentoring, and continuous reliability improvement.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →