9 дней назад
Senior Site Reliability Engineer (AI Infrastructure)
187 040 - 359 720$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI Infrastructure): Building and operating globally distributed, fault-tolerant infrastructure for TikTok’s Recommendation and Search engines with an accent on production ownership, observability, automation, and service reliability. Focus on designing distributed systems, leading infrastructure migrations, managing capacity and SLAs/SLOs, and solving systemic causes of complex production incidents.
Location: San Jose, United States; fully in-person schedule up to 5 days a week
Salary: $187,040–$359,720 annually, with potential additional bonuses, incentives, and restricted stock units.
Company
operates a secure U.S. data privacy and cybersecurity environment for TikTok apps, content, algorithms, and the U.S. user ecosystem.
What you will do
- Partner with engineering and product teams across system design, architecture reviews, deployment, operations, and continuous service refinement.
- Build tools, platforms, and automation that improve reliability, scalability, R&D efficiency, and operational workflows.
- Monitor service health, latency, and key metrics to maintain high availability and performance across large-scale, multi-region systems.
- Lead infrastructure migrations and architecture upgrades in a highly restricted compliance environment.
- Manage capacity, resource allocation, stability optimization, error attribution, and SLA/SLO compliance.
- Drive incident management, blameless postmortems, and systemic root-cause resolution.
Requirements
- Bachelor’s degree or equivalent practical experience in computer science, software engineering, or a related technical field.
- 3+ years of hands-on experience in SRE, DevOps, or systems engineering for large-scale, highly reliable systems.
- Strong knowledge of Linux, system performance, networking fundamentals, and distributed systems troubleshooting.
- Programming experience in at least one of Go, Python, C/C++, or Java, plus Bash or shell scripting.
- Experience with CI/CD, automated deployment pipelines, complex production environments, and operational workflow modernization.
- Ability to work on-site in San Jose in a fully in-person schedule of up to five days per week.
Nice to have
- Experience with Kubernetes, service mesh architectures, cloud platforms, and Infrastructure-as-Code such as Terraform.
- Experience with Prometheus, Grafana, distributed tracing, or other observability tooling.
- Experience building self-service automation, distributed systems, or infrastructure.
- Interest or experience using LLMs, agentic AI, or machine learning concepts for operational automation and troubleshooting.
Culture & Benefits
- Medical, dental, and vision insurance from day one, plus a 401(k) savings plan with company match.
- Paid parental leave, short- and long-term disability coverage, life insurance, and wellbeing benefits.
- 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with accrual increasing by tenure.
- Work in a secure, isolated infrastructure environment during a critical 0-to-1 platform-building phase.
- Inclusive workplace with reasonable accommodations available during recruitment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
9 дней назад
Site Reliability Engineer, Edge Services - USDS
136 800 - 359 720$
9 дней назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
6 дней назад
Site Reliability Engineer (AWS)
180 000 - 220 000$
Okta
7 дней назад
Senior TDI Site Reliability Engineer (AWS)
165 000 - 225 600$
Okta
7 дней назад
Senior TDI Site Reliability Engineer, Okta Federal
165 000 - 225 600$