Назад
Company hidden
5 дней назад

Staff Site Reliability Engineer (Cybersecurity)

199 750 - 270 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (Python/Terraform, AWS/Kubernetes): Establishing and evolving reliability strategy, observability standards, incident response, and service ownership for a cybersecurity platform with an accent on distributed systems, SLIs/SLOs, error budgets, and production operations. Focus on leading cross-functional reliability initiatives, improving on-call and recovery readiness, and building sustainable engineering-wide SRE practices.

Location: US, Remote; up to 10% travel for team off-sites and in-person project kick-offs.

Base salary: $199,750–$270,000 annually, plus equity eligibility.

Company

hirify.global is a fast-growing cybersecurity company developing NodeZero, a platform for production-safe autonomous penetration testing and security assessment operations.

What you will do

  • Own and evolve the engineering-wide SRE strategy, operating model, and reliability standards.
  • Coordinate Infrastructure, product, service, security, and business stakeholders on reliability, observability, incident response, and operational readiness.
  • Establish service ownership, meaningful SLIs and SLOs, error budgets, dashboards, actionable alerts, runbooks, and escalation paths.
  • Lead complex cross-functional reliability initiatives and set standards for incident command, on-call health, post-incident learning, and recovery readiness.
  • Shape the technical direction and growth path of the SRE function while participating in a 24/7 on-call rotation.

Requirements

  • Experience designing, operating, and troubleshooting large-scale distributed systems in production.
  • Deep knowledge of reliability engineering, observability, incident management, and production operations.
  • Experience establishing SLIs, SLOs, actionable alerts, observability, and service ownership.
  • Backend experience building automation that reduces toil, strengthens safeguards, and improves operational efficiency.
  • Experience leading high-severity incidents and improving incident response programs.
  • Python and Terraform, or equivalent automation and infrastructure-as-code tools; production experience with AWS, Kubernetes, observability platforms, and CI/CD or GitOps workflows.

Culture & Benefits

  • Fully remote work with a collaborative, inclusive, and ownership-oriented culture.
  • Health, vision, and dental insurance for employees and families.
  • Flexible vacation policy and generous parental leave.
  • Career development opportunities, competitive compensation, and stock-option eligibility.
  • Regular team off-sites and in-person project kick-offs may require limited travel.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →