13 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
122 574 - 259 200$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM): Managing data services and real-time/batch pipelines with an accent on reliability engineering, observability, incident response, and AI-powered automation. Focus on designing LLM-powered operational tools, tuning distributed systems, and maintaining 24-hour service coverage through on-call rotations.
Location: San Jose, United States; fully in-person schedule up to 5 days a week
Salary: $122,574–$259,200 annually, with potential discretionary bonuses, incentives, and restricted stock units.
Company
USDS is a TikTok joint venture focused on data privacy, cybersecurity, national security, trust and safety, and protecting U.S. user data.
What you will do
- Manage daily operations for data services and real-time or batch pipelines, including SLA, SLO, and SLI management.
- Deploy systems, tune performance, troubleshoot incidents, and improve service reliability.
- Design and deploy AI agents and LLM-powered automation for incident response, root cause analysis, and proactive monitoring.
- Build tools and automation to improve system administration and operational efficiency.
- Support the full service lifecycle from design and capacity planning through launch, deployment, operation, and refinement.
- Participate in on-call rotations providing 24-hour coverage and conduct incident response and postmortems.
Requirements
- Bachelor’s degree or higher in computer science or a related technical discipline.
- At least 1 year of industrial experience.
- Experience integrating AI or LLM APIs into internal workflows or infrastructure tooling.
- Strong independent thinking, troubleshooting, and problem-solving skills.
- Knowledge of Unix/Linux internals, networking, distributed systems, monitoring, and observability.
- Experience with monitoring tools such as Prometheus, Grafana, or DataDog.
Nice to have
- Advanced knowledge of Unix/Linux systems, networking fundamentals, and system performance tuning.
- Familiarity with MySQL, Redis, Nginx, Kafka, Kubernetes, Docker, Hadoop, Spark, Flink, Hive, OLAP, or ClickHouse.
Culture & Benefits
- Medical, dental, and vision insurance from day one.
- 401(k) savings plan with company match.
- Paid parental leave, disability coverage, life insurance, and wellbeing benefits.
- 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
- Inclusive workplace with reasonable accommodations available during recruitment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
19 часов назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
6 дней назад
Staff Engineer (Platform)
175 000 - 250 000$
21 час назад
Platform Engineer (AI)
140 000 - 180 000$
20 часов назад
Staff Platform Engineer (AI/ML)
140 800 - 176 000$
6 дней назад
Senior Platform Engineer (Build & Developer Infrastructure)
150 000 - 200 000$