Назад
Company hidden
7 дней назад

Staff Software Engineer, Reliability

203 500 - 248 500$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Software Engineer, Reliability (AWS/Java/Scala): Building reliability systems for mission-critical mobility infrastructure with an accent on multi-region failover, observability, database replication, and resilience engineering. Focus on designing disaster recovery and dependency-failover systems, implementing chaos engineering and SLO-based monitoring, and reducing incident recovery time while maintaining 99.9%+ uptime.

Location: Seattle, Washington, United States; on-site at least four days per week

Base salary: $203,500–$248,500 USD annually

Company

hirify.global builds AI-powered mobility commerce and recognition technologies for parking, retail, hospitality, and real estate.

What you will do

  • Own the reliability posture of the platform, defining reliability practices, metrics, and systems to maintain 99.9%+ uptime.
  • Design automatic failover, circuit breakers, retry policies, degraded modes, and local mirrors for critical external dependencies such as Twilio and Stripe.
  • Architect active-passive or active-active multi-region deployments with database replication, automated failover, DNS traffic routing, and disaster recovery testing.
  • Build observability with Datadog, including APM, logs, metrics correlation, synthetic monitoring, SLO alerting, and customer-impact dashboards.
  • Lead incident management, on-call and escalation processes, post-mortems, runbook automation, and MTTR reduction initiatives.
  • Drive adoption of resilience patterns such as health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering.

Requirements

  • 8+ years of software engineering, reliability engineering, SRE, or large-scale production operations experience.
  • Expertise in multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery.
  • Production observability experience with monitoring, alerting, tracing, and logging systems, specifically Datadog or a similar APM platform.
  • Strong distributed-systems and database expertise, including replication, backup and restore, connection pooling, query optimization, and relational and NoSQL databases.
  • Production AWS experience with multi-region deployments, load balancing, and DNS-based failover.
  • Expert-level Java and/or Scala proficiency, including JVM performance, concurrency, and operational characteristics.
  • Ability to work from the Seattle office at least four days per week.

Nice to have

  • Scala experience and reliability engineering experience at operationally excellent companies or high-growth startups.
  • Incident response leadership, blameless post-mortems, and MTTR reduction experience.
  • Chaos engineering with tools such as Chaos Monkey or Gremlin, including game days and failure-injection testing.
  • Hyperscale performance optimization, profiling, benchmarking, capacity planning, and system tuning.
  • Open-source contributions or technical writing in reliability engineering, distributed systems, or production operations.

Culture & Benefits

  • Office-first collaboration model focused on in-person interaction and innovation.
  • Inclusive environment where employees are encouraged to contribute ideas.
  • Healthcare benefits, 401(k), disability coverage, life insurance, stock options, and bonus plans may be available.
  • AI-assisted development tools include GitHub Copilot.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →