19 часов назад
Site Reliability Engineer, Reliability Team - USDS
122 574 - 259 200$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Reliability Team - USDS (SRE/Distributed Systems): Building and operating large-scale, fault-tolerant production systems with an accent on automation, observability, disaster recovery, and incident response. Focus on designing multi-region failover, coordinating high-priority incident restoration, analyzing root causes, and managing capacity for massive traffic surges.
Location: San Jose, United States; fully in-person schedule up to 5 days a week
Salary: $122,574–$259,200 annually
Company
USDS is a TikTok joint venture focused on data privacy, cybersecurity, national security, and protecting U.S. user data, applications, and algorithms.
What you will do
- Design, optimize, and operate high-concurrency distributed systems for scalability, reliability, and high availability.
- Build automation tools, streamline deployments, and manage infrastructure as code.
- Develop monitoring, alerting, logging, and SLI/SLO systems to improve service observability.
- Design and run global disaster recovery drills, simulate failures, and validate multi-region failover mechanisms.
- Respond to high-priority production incidents, coordinate cross-functional war rooms, and drive service restoration.
- Lead blameless post-mortems, root-cause analysis, continuous improvement, and capacity planning.
Requirements
- Bachelor’s degree in Computer Science, a related technical field, or equivalent practical experience.
- Proficiency in one or more programming languages, including Go, Python, Java, or C++.
- Strong knowledge of Linux system internals, networking fundamentals including TCP/IP and DNS, load balancing, and distributed systems.
- Experience managing containerized environments such as Kubernetes or Docker.
- Experience in high-traffic production environments, incident response, site stability, or disaster recovery is preferred.
- Experience with observability tools, infrastructure as code, multi-region failover, and distributed database consistency is preferred.
Culture & Benefits
- On-site collaboration supports rapid decision-making, team development, and integrated execution.
- Medical, dental, and vision insurance from day one, plus a 401(k) savings plan with company match.
- Paid parental leave, disability coverage, life insurance, and wellbeing benefits.
- 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with increasing accruals by tenure.
- Inclusive workplace with reasonable accommodations available during recruitment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
19 часов назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
7 часов назад
Site Reliability Engineer, Edge Services - USDS
136 800 - 359 720$
9 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
4 дня назад
Reliability Engineer (SRE)
75 000 - 95 000$
3 дня назад
Systems Reliability Engineer (Kubernetes/Cloud)
100 000 - 150 000$
5 дней назад
Sr Software Development Engineer (US Federal)
163 800 - 245 800$