обновлено 5 дней назад
Staff Site Reliability Engineer, Federal (TS/SCI)
174 000 - 238 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer, Federal (TS/SCI) (Cloud Infrastructure and Federal Compliance): Building and operating reliable, scalable, and secure cloud services for federal customers with an accent on Kubernetes, infrastructure automation, observability, and FedRAMP/IL6 compliance. Focus on leading incident response, eliminating operational toil with Go, Python, and Terraform, and improving production reliability across complex distributed systems.
Location: Washington, DC; employee must be on U.S. soil and maintain or obtain a U.S. security clearance
Annual base salary: $174,000–$238,000 USD, with potential equity, bonus, and benefits.
Company
Okta builds identity and security infrastructure that helps organizations securely operate cloud services and adopt AI-enabled technologies.
What you will do
- Design, build, and operate large-scale cloud infrastructure and customer-facing production services.
- Lead incident response, post-incident reviews, and initiatives to improve availability, scalability, performance, and resilience.
- Define and improve SLIs, SLOs, error budgets, observability, monitoring, and production telemetry.
- Develop automation, infrastructure, internal platforms, operational guardrails, and self-service tooling using Go, Python, Terraform, and related technologies.
- Improve deployment safety and operational workflows through CI/CD, GitOps, and platform engineering.
- Lead cross-team reliability initiatives, mentor engineers, influence architecture, and own projects through production rollout.
Requirements
- Active U.S. TS/SCI clearance with Full Scope Poly is required.
- Must be on U.S. soil and able to obtain and maintain a U.S. security clearance as required by federal contracts.
- Strong experience operating large-scale production services in AWS and/or GCP and deep production Kubernetes expertise.
- Advanced experience with Kubernetes troubleshooting, Terraform, Helm, cloud networking, observability, and infrastructure security.
- Strong software engineering skills in Go and/or Python, plus experience building automation and internal engineering platforms.
- Experience with incident response, reliability engineering, CI/CD, distributed data platforms, and leading initiatives across multiple teams.
Nice to have
- Experience with FedRAMP and Impact Level 6 (IL6) compliance frameworks.
- Experience operating SaaS platforms, Kubernetes-based microservices, or globally distributed production environments.
- Experience with GitOps, ArgoCD, and AI-assisted operational tooling.
Culture & Benefits
- Automation-first environment focused on operational excellence, reliability, security, and engineering velocity.
- Collaboration with distributed engineering organizations across multiple time zones and cultures.
- Health, dental, and vision insurance, 401(k), flexible spending account, PTO, and parental leave.
- Equity and bonus opportunities where applicable.
- In-person onboarding designed to support connection and accelerate impact.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →