6 дней назад
Senior Site Reliability Engineer (SRE)
187 040 - 359 720$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (SRE): Building and operating large-scale, fault-tolerant production systems with an accent on automation, observability, disaster recovery, and incident response. Focus on designing multi-region failover, leading complex recovery drills, resolving high-priority outages, and improving reliability through root-cause analysis.
Location: San Jose R&D; fully in-person up to 5 days a week
Salary: $187,040–$359,720 annually, plus potential discretionary bonuses, incentives, and restricted stock units.
Company
operates TikTok apps and focuses on data privacy, cybersecurity, national security, and protection of U.S. user data.
What you will do
- Design, optimize, and operate high-concurrency distributed systems for scalability, reliability, and high availability.
- Build automation tools, streamline deployments, and manage infrastructure as code.
- Develop monitoring, alerting, logging, and SLI/SLO systems to improve service observability.
- Design and lead global disaster recovery drills, including complex failure simulations and failover validation.
- Respond to high-priority production incidents, coordinate cross-functional war rooms, and drive service restoration.
- Conduct blameless post-mortems, root-cause analysis, capacity planning, and reliability improvements.
Requirements
- Bachelor’s degree in Computer Science or a related technical field, or equivalent practical experience.
- Proficiency in one or more programming languages, such as Go, Python, Java, or C++.
- Strong knowledge of Linux internals, networking including TCP/IP and DNS, load balancing, and distributed systems.
- Experience managing containerized environments such as Kubernetes and Docker.
- Experience in high-traffic production environments, incident response, disaster recovery, multi-region failover, or distributed database consistency.
- Familiarity with observability tools and infrastructure as code.
Culture & Benefits
- On-site collaboration focused on speed, alignment, and integrated execution.
- Medical, dental, and vision insurance from day one.
- 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
- Wellbeing benefits, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
- Inclusive workplace with reasonable accommodations available during recruitment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
6 дней назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
2 дня назад
Senior Site Reliability Engineering
182 800 - 247 300$
Reddit
5 дней назад
Staff Site Reliability Engineer, Ads
217 000 - 303 900$
6 дней назад
Site Reliability Engineer, Edge Services - USDS
136 800 - 359 720$
2 дня назад
DevOps & SRE Engineer (Kubernetes)
100 000 - 150 000$