Назад
обновлено 2 дня назад

Senior Software Engineer (Reliability)

196 000 - 230 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Software Engineer (Reliability): Building incident response, observability, and reliability tooling for Robinhood's production systems with an accent on incident leadership, fault-tolerant architecture, and multi-region operations. Focus on designing monitoring frameworks, coordinating high-severity mitigation, and driving measurable improvements in MTTD, MTTR, availability, and customer impact.

Location: Menlo Park, California office with in-person attendance expected at least 3 days per week. Base pay locations also include New York, Bellevue, and Washington, DC.

Salary: $196,000–$230,000 USD annual base pay, plus bonus opportunities, equity, and benefits.

Company

Robinhood is a fintech company building products and infrastructure to democratize finance.

What you will do

  • Help establish the Robinhood Command Center as a front line for detecting, coordinating, and mitigating production incidents.
  • Lead incident mitigation by coordinating service owners, facilitating rollbacks and traffic shifts, and maintaining a clear source of truth.
  • Develop incident management processes, response tooling, education, adoption, and metrics for MTTD and MTTR improvements.
  • Define company-wide dashboards, alerts, and observability frameworks for critical user journeys, availability, and business-impact metrics.
  • Design failure-mitigation strategies, monitoring standards, and observability roadmaps across multi-region and multi-cluster systems.
  • Drive post-incident governance, postmortems, SEV reviews, executive reporting, mentoring, and engineering culture initiatives.

Requirements

  • 5+ years of software engineering experience, including significant experience operating production systems.
  • 2+ years focused on reliability engineering, infrastructure, distributed systems, or production operations.
  • Hands-on experience as an incident commander, primary on-call engineer, or in a similar incident leadership role.
  • Deep knowledge of systems reliability, observability frameworks, fault-tolerant architecture, capacity planning, and failover strategies.
  • Experience with multi-region or multi-cluster architectures and modern observability stacks such as OpenTelemetry, Prometheus, and Grafana.
  • Strong communication and cross-functional collaboration skills, with a track record of improving MTTD, MTTR, availability, or customer impact.

Culture & Benefits

  • High-impact work on reliability and infrastructure for financial products.
  • Performance-driven compensation with bonus programs, equity ownership, and 401(k) matching.
  • Health insurance, with 100% employee coverage and 90% dependent coverage.
  • Flexible lifestyle wallet for wellness, learning, and other expenses.
  • Employer-paid life and disability insurance, fertility benefits, mental health benefits, paid time off, sick time, parental leave, and company holidays.
  • Office experience with catered meals, events, and comfortable workspaces.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →