3 дня назад
Senior Incident Commander (SRE/Python)
187 000 - 233 500$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Incident Commander (SRE/Python): Leading critical production incident response and building automation and tooling for Box’s cloud operations with an accent on reliability, observability, and rapid service recovery. Focus on coordinating cross-functional incident bridges, reducing MTTx, improving change and incident management, and strengthening service resiliency across globally diverse 24x7 environments.
Location: Redwood City, California, United States; hybrid work with at least 3 days per week from the assigned office
Base salary: $187,000–$233,500 USD per year, plus equity and benefits eligibility.
Company
provides an intelligent content management platform for collaboration, content lifecycle management, security, enterprise AI, and business workflow transformation.
What you will do
- Own and direct critical, blocker, and other high-severity production incidents from escalation through mitigation and recovery.
- Lead incident bridges, clarify customer impact, coordinate subject-matter experts, and facilitate cross-functional response.
- Improve incident platform tooling, automation, templates, runbooks, and measurable response workflows using Python, APIs, PagerDuty, Jira, and internal services.
- Partner with SRE and engineering teams to improve observability, operability, dependency knowledge, service manageability, and resiliency.
- Lead daily change reviews and evaluate production change risk, including emergency and break-glass changes.
- Convert incident learnings into engineering improvements such as SLOs, signal-quality enhancements, rollback readiness, dependency documentation, and reliability projects.
Requirements
- 5+ years of experience in SRE, production operations, reliability engineering, or high-scale internet/SaaS operations, including major production incident leadership.
- Strong Incident Commander or Technical Duty Officer skills, including delegation, escalation judgment, clear communication, and calm decision-making under uncertainty.
- Practical expertise with SLIs/SLOs, error budgets, observability, golden signals, blameless postmortems, toil reduction, and product-team reliability partnerships.
- Proficiency in Python for maintainable automation and tooling, including APIs, packaging or service-style tools, and code review.
- Experience troubleshooting Linux/Unix and distributed systems, with networking knowledge covering DNS, TLS, load balancing, HTTP, routing, and firewalls.
- Experience with cloud environments, preferably GCP, plus Kubernetes or equivalent container and orchestration concepts; strong written and verbal communication and mentoring skills.
Nice to have
- Experience in a 24x7 NOC, GTOC, or follow-the-sun operations center.
- Experience with Prometheus-compatible observability, distributed tracing, synthetic monitoring, PagerDuty, ChatOps, workflow automation, and incident drills.
- Additional operations automation tools such as Go, shell, Terraform, or CI/CD.
- Experience with dependency graphs, service catalogs, Tier 1 journey documentation, or change management under pressure.
Culture & Benefits
- In-person collaboration and community are core parts of the working culture.
- Work takes place in a globally diverse, 24x7 operational environment.
- The role is eligible for equity and employee benefits.
- supports diversity, inclusion, and reasonable accommodations throughout the application and interview process.
Hiring process
- The recruiter provides additional details about the working model and company culture during the hiring process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Datadog
4 дня назад
Senior Software Engineer - Incident Insights & Readiness (SRE)
192 000 - 240 000$
4 дня назад
Staff Service Reliability and Operational Intelligence Engineer (AI Ops)
152 000 - 228 000$
6 дней назад
Senior Engineer (SRE/Incident Management)
100 000 - 215 000$
6 дней назад
Senior Staff Engineer - SRE (Incident Prevention)
110 000 - 260 000$
6 дней назад
Senior Staff Engineer (SRE)
120 000 - 260 000$
6 дней назад
Staff Engineer (SRE)
110 000 - 230 000$