Назад
обновлено 5 дней назад

Staff Site Reliability Engineer, Ads

217 000 - 303 900$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer, Ads (Distributed Systems/SRE): Leading reliability engineering for Reddit’s critical user-facing systems across APIs, feeds, content delivery, search, messaging, and real-time experiences with an accent on availability, latency, scalability, and operational excellence. Focus on architecting highly available systems, automating incident response and remediation, reducing operational risk, and resolving complex reliability challenges at internet scale.

Location: San Francisco, CA, United States

Salary: $217,000–$303,900 USD annually, plus potential equity and commission depending on the position.

Company

Reddit is a large-scale community platform with more than 100,000 active communities and approximately 130 million daily active unique visitors.

What you will do

  • Lead reliability, scalability, performance, and operational excellence for critical user-facing systems and services.
  • Partner with product and infrastructure teams on highly available architectures, failover, redundancy, graceful degradation, traffic management, and capacity planning.
  • Identify systemic risks and reliability bottlenecks across services, dependencies, deployments, and infrastructure, then drive long-term mitigation.
  • Build automation and tooling for safer deployments, incident response, remediation workflows, and reliability guardrails.
  • Lead complex incident response, blameless postmortems, root-cause analysis, and sustainable fixes.
  • Define reliability standards around SLIs, SLOs, capacity management, release engineering, and operational maturity while mentoring engineers.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
  • Experience supporting high-traffic, user-facing production environments and designing highly available systems.
  • Deep knowledge of distributed systems, networking, Linux systems, or cloud-native architectures.
  • Strong programming skills in Go, Python, or similar languages.
  • Understanding of observability systems, including metrics, logging, tracing, and alerting.
  • Experience with SLOs, automation, incident management, performance optimization, and troubleshooting across applications, infrastructure, networking, and services.

Nice to have

  • Experience operating systems at internet-scale traffic volumes and leading large-scale incident response or operational transformation.
  • Experience with Kubernetes, containers, cloud infrastructure, and modern deployment platforms.
  • Familiarity with Prometheus, Grafana, OpenTelemetry, Envoy, Kafka, ClickHouse, Cassandra, Redis, or similar distributed infrastructure technologies.
  • Experience with CDN optimization, edge reliability, traffic engineering, or global infrastructure.
  • Open-source contributions or participation in technical communities.

Culture & Benefits

  • Global benefits covering workspace, professional development, caregiving, and family planning.
  • Gender-affirming care, mental health and coaching benefits, and private medical, dental, and vision benefits.
  • Retirement savings account with matching contribution.
  • Flexible vacation, paid volunteer time off, and generous paid parental leave.
  • Opportunity to influence reliability and performance for a consumer platform used by millions of people daily.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →