8 дней назад
Site Reliability Engineer (Defense Systems)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Defense Systems): Owns the reliability, performance, observability, and operational health of critical engineering systems supporting software development, CI/CD, testing, and developer workflows with an accent on Linux, networking, cloud, Kubernetes, and incident response. Focus on investigating complex cross-system failures, improving monitoring and capacity, and driving evidence-backed corrective actions for mission-critical services.
Location: Allen, Texas, or Torrance, California. The position may require access to classified information or restricted U.S. Government sites, systems, or information, subject to applicable eligibility and authorization requirements.
Company
develops and delivers advanced defense systems.
What you will do
- Own reliability, availability, latency, capacity, recovery expectations, and operational health for critical engineering services.
- Investigate and resolve complex incidents across applications, Linux, networking, storage, Kubernetes, cloud, and other system layers.
- Build and improve monitoring, observability, alerting, and diagnostic systems.
- Analyze performance and capacity across compute, memory, storage, networking, connections, and other constrained resources.
- Lead root cause analyses, postmortems, and corrective and preventive actions.
- Partner with DevOps, Cloud, Software, Security, Test, and IT teams and participate in the on-call rotation.
Requirements
- Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or a related technical field.
- 5+ years of experience in Site Reliability Engineering, Production Engineering, Systems Engineering, Infrastructure Engineering, or a related discipline supporting production or mission-critical systems.
- Strong Linux expertise covering CPU, memory, storage, networking, processes, sockets, system services, and performance and capacity concepts.
- Experience operating observability, monitoring, alerting, and incident response systems.
- Strong networking and application fundamentals, including TCP, TLS, HTTP, DNS, reverse proxies, load balancers, connection states, and timeouts.
- Ability to satisfy applicable U.S. Government security clearance, authorization, citizenship, or other eligibility requirements if designated for the position. Access to export-controlled information requires eligibility without additional export licensing.
Culture & Benefits
- Significant ownership and direct impact in an early-stage organization.
- Cross-functional collaboration across engineering and infrastructure disciplines.
- Generous benefits package.
- Equal opportunity employment and reasonable accommodations throughout the hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Vapi
12 дней назад
Senior Site Reliability Engineer (AI)
280 000 - 314 000$
9 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
14 дней назад
Senior Site Reliability Engineer (Temporal)
180 000 - 200 000$
11 дней назад
Site Reliability Engineer - Vice President (Kubernetes)
130 000 - 160 000$
9 дней назад
Site Reliability Engineer (Kubernetes)
180 000 - 220 000$
9 дней назад