8 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM): Managing data services and real-time and batch pipelines for machine-learning systems that detect internet abuse and fraud with an accent on reliability, observability, incident response, and AI-powered automation. Focus on deploying LLM-powered incident tooling, tuning distributed systems, troubleshooting production services, and supporting scheduled 24/7 operations.
Location: Seattle, United States; fully in-person schedule up to 5 days a week
Salary: $129,960–$246,240 annually, with potential discretionary bonuses, incentives, and restricted stock units.
Company
Joint Venture operates data privacy, cybersecurity, trust and safety, and content moderation systems for TikTok's U.S. platform and users.
What you will do
- Manage day-to-day operations of data services and real-time and batch data pipelines, including SLA, SLO, and SLI management.
- Deploy AI agents and LLM-powered automation for incident response, root cause analysis, and proactive monitoring.
- Create administration tools and automation to improve operational efficiency and delivery quality.
- Support the full service lifecycle, from design and capacity planning through launch, deployment, operation, and refinement.
- Handle user support, incident response, and postmortems.
- Participate in a 24/7 support rotation with scheduled shifts that may include holidays.
Requirements
- Bachelor's degree or higher in computer science or a related technical discipline.
- At least 1 year of industrial experience.
- Experience integrating AI or LLM APIs into internal workflows or infrastructure tooling.
- Independent thinking and strong troubleshooting skills.
- Familiarity with Unix/Linux internals, networking, distributed systems, monitoring tools, and observability practices.
- Experience with Prometheus, Grafana, DataDog, or comparable monitoring platforms; backend and big data technologies such as MySQL, Redis, Nginx, Kafka, Kubernetes, Docker, Hadoop, Spark, Flink, Hive, OLAP, or ClickHouse is relevant.
Culture & Benefits
- On-site work supports fast alignment, real-time decision-making, team development, and integrated execution.
- Medical, dental, and vision insurance from day one.
- 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
- Wellbeing benefits, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
- Inclusive workplace with reasonable accommodations available during recruitment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
12 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
122 574 - 259 200$
1 день назад
Senior Site Reliability Engineer (AI/LLM)
187 040 - 359 720$
22 часа назад
Software Engineer, AI Infrastructure - USDS
129 960 - 246 240$
24 часа назад
Site Reliability Engineer (Tech Infrastructure)
129 960 - 246 240$
17 часов назад
Senior Site Reliability Engineer (Compute)
177 688 - 341 734$
21 час назад
Site Reliability Engineer, Product - USDS
122 574 - 259 200$