5 дней назад
Staff Site Reliability Engineer (Cybersecurity)
199 750 - 270 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Site Reliability Engineer (Python/Terraform, AWS/Kubernetes): Establishing and evolving reliability strategy, observability standards, incident response, and service ownership for a cybersecurity platform with an accent on distributed systems, SLIs/SLOs, error budgets, and production operations. Focus on leading cross-functional reliability initiatives, improving on-call and recovery readiness, and building sustainable engineering-wide SRE practices.
Location: US, Remote; up to 10% travel for team off-sites and in-person project kick-offs.
Base salary: $199,750–$270,000 annually, plus equity eligibility.
Company
is a fast-growing cybersecurity company developing NodeZero, a platform for production-safe autonomous penetration testing and security assessment operations.
What you will do
- Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards.
- Coordinate Infrastructure, product, service, security, and business stakeholders on reliability, observability, incident response, and operational readiness.
- Establish service ownership, meaningful SLIs and SLOs, error budgets, dashboards, actionable alerts, runbooks, and escalation paths.
- Lead complex cross-functional reliability initiatives and set standards for incident command, on-call health, post-incident learning, and recovery readiness.
- Shape the technical direction and growth path of the SRE function while participating in a 24/7 on-call rotation.
Requirements
- Experience designing, operating, and troubleshooting large-scale distributed systems in production.
- Deep knowledge of reliability engineering, observability, incident management, and production operations.
- Experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
- Backend experience building automation that reduces toil, strengthens safeguards, and improves operational efficiency.
- Experience leading high-severity incidents and improving incident response programs.
- Python and Terraform, or equivalent automation and infrastructure-as-code tools; production experience with AWS, Kubernetes, observability platforms, and CI/CD or GitOps workflows.
Culture & Benefits
- Fully remote work with a collaborative, inclusive, and ownership-oriented culture.
- Health, vision, and dental insurance for employees and families.
- Flexible vacation policy and generous parental leave.
- Career development opportunities, competitive compensation, and stock-option eligibility.
- Regular team off-sites and in-person project kick-offs may require limited travel.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 дней назад
Lead Site Reliability Engineer (Cybersecurity)
145 000 - 200 000$
6 дней назад
Principal Site Reliability Engineer (Kubernetes)
190 000 - 220 000$
12 дней назад
Site Reliability Engineering Manager (AWS/Kubernetes)
205 000 - 255 000$
9 дней назад
Cloud Site Reliability Engineer (AWS)
120 000 - 130 000$
8 дней назад
Senior Site Reliability Engineer (Healthcare)
200 000 - 240 000$
12 дней назад
Site Reliability Engineer, Observability
160 000 - 200 000$