Назад
Company hidden
19 часов назад

Site Reliability Engineer, Reliability Team - USDS

122 574 - 259 200$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer, Reliability Team - USDS (SRE/Distributed Systems): Building and operating large-scale, fault-tolerant production systems with an accent on automation, observability, disaster recovery, and incident response. Focus on designing multi-region failover, coordinating high-priority incident restoration, analyzing root causes, and managing capacity for massive traffic surges.

Location: San Jose, United States; fully in-person schedule up to 5 days a week

Salary: $122,574–$259,200 annually

Company

USDS is a TikTok joint venture focused on data privacy, cybersecurity, national security, and protecting U.S. user data, applications, and algorithms.

What you will do

  • Design, optimize, and operate high-concurrency distributed systems for scalability, reliability, and high availability.
  • Build automation tools, streamline deployments, and manage infrastructure as code.
  • Develop monitoring, alerting, logging, and SLI/SLO systems to improve service observability.
  • Design and run global disaster recovery drills, simulate failures, and validate multi-region failover mechanisms.
  • Respond to high-priority production incidents, coordinate cross-functional war rooms, and drive service restoration.
  • Lead blameless post-mortems, root-cause analysis, continuous improvement, and capacity planning.

Requirements

  • Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
  • Proficiency in one or more programming languages, including Go, Python, Java, or C++.
  • Strong knowledge of Linux system internals, networking fundamentals including TCP/IP and DNS, load balancing, and distributed systems.
  • Experience managing containerized environments such as Kubernetes or Docker.
  • Experience in high-traffic production environments, incident response, site stability, or disaster recovery is preferred.
  • Experience with observability tools, infrastructure as code, multi-region failover, and distributed database consistency is preferred.

Culture & Benefits

  • On-site collaboration supports rapid decision-making, team development, and integrated execution.
  • Medical, dental, and vision insurance from day one, plus a 401(k) savings plan with company match.
  • Paid parental leave, disability coverage, life insurance, and wellbeing benefits.
  • 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with increasing accruals by tenure.
  • Inclusive workplace with reasonable accommodations available during recruitment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →