2 часа назад
Site Reliability Engineer
170 000 - 195 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Kubernetes/AWS): Building and operating reliable cloud deployments and edge-device infrastructure for defense technology with an accent on observability, incident response, and infrastructure as code. Focus on defining SLIs and SLOs, debugging Kubernetes and stateful workloads, and hardening systems across remote edge environments.
Location: El Segundo, California, United States. Applicants must be U.S. citizens or nationals, lawful permanent residents, refugees, asylees, or eligible to obtain the required U.S. Government authorizations.
Base salary: $170,000–$195,000 per year, plus equity and bonus opportunities.
Company
is a venture-backed defense technology startup building infrastructure that unifies sensors, autonomy, and operators for national security operations.
What you will do
- Own production reliability across cloud deployments and edge devices operating in remote and contested environments.
- Define and drive reliability SLIs and SLOs, including error-budget practices.
- Manage the observability stack with Grafana, Prometheus, Loki, and OpenTelemetry.
- Participate in on-call rotations, incident response, evidence-first troubleshooting, and blameless postmortems.
- Encode reliability and security practices in infrastructure as code.
- Help establish infrastructure and DevSecOps practices in a fast-moving startup environment.
Requirements
- 3+ years of experience as an SRE or in a related role.
- Deep Kubernetes operations experience, including node lifecycle, workload scheduling, StatefulSets, graceful drains, and live cluster debugging.
- Experience designing observability dashboards and high-signal alerting rules.
- Production experience with Terraform or OpenTofu and fluent AWS knowledge covering IAM, networking, multi-account environments, and workload hardening.
- Experience managing highly available database deployments and IoT or edge device fleets.
- Eligibility to work on projects subject to U.S. Government export regulations is required.
Nice to have
- Experience with GovCloud, FIPS, regulated, or air-gapped environments.
- Experience operating constrained edge hardware such as NVIDIA Jetson platforms.
- Overlay or mesh networking experience with Nebula, WireGuard, Tailscale, or similar tools.
- Experience setting up SLO and error-budget tooling from scratch.
- Active security clearance.
Culture & Benefits
- Significant stock options as an early-stage company.
- 401(k) with employer matching and full medical, dental, and vision coverage.
- Relocation assistance may be provided.
- Unlimited PTO with a two-week minimum, plus 11 paid holidays.
- Paid parental leave for both parents.
- Office perks include lunch, a stocked kitchenette, free EV charging, and coffee, snacks, and craft beer.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 часов назад
Site Reliability Engineer (TypeScript)
180 000 - 220 000$
13 часов назад
Site Reliability Engineer (Aerospace)
18 часов назад
Principal Site Reliability Engineer (AWS/Kubernetes)
163 620 - 212 710$
10 часов назад
Senior DevOps Engineer (AWS/Kubernetes)
185 000 - 218 000$
12 часов назад
Platform Engineer
200 000 - 300 000$
11 часов назад
Senior Site Reliability Engineer (AWS)
180 000 - 200 000$