Назад
Company hidden
13 часов назад

Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)

122 574 - 259 200$
Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM): Managing data services and real-time/batch pipelines with an accent on reliability engineering, observability, incident response, and AI-powered automation. Focus on designing LLM-powered operational tools, tuning distributed systems, and maintaining 24-hour service coverage through on-call rotations.

Location: San Jose, United States; fully in-person schedule up to 5 days a week

Salary: $122,574–$259,200 annually, with potential discretionary bonuses, incentives, and restricted stock units.

Company

USDS is a TikTok joint venture focused on data privacy, cybersecurity, national security, trust and safety, and protecting U.S. user data.

What you will do

  • Manage daily operations for data services and real-time or batch pipelines, including SLA, SLO, and SLI management.
  • Deploy systems, tune performance, troubleshoot incidents, and improve service reliability.
  • Design and deploy AI agents and LLM-powered automation for incident response, root cause analysis, and proactive monitoring.
  • Build tools and automation to improve system administration and operational efficiency.
  • Support the full service lifecycle from design and capacity planning through launch, deployment, operation, and refinement.
  • Participate in on-call rotations providing 24-hour coverage and conduct incident response and postmortems.

Requirements

  • Bachelor’s degree or higher in computer science or a related technical discipline.
  • At least 1 year of industrial experience.
  • Experience integrating AI or LLM APIs into internal workflows or infrastructure tooling.
  • Strong independent thinking, troubleshooting, and problem-solving skills.
  • Knowledge of Unix/Linux internals, networking, distributed systems, monitoring, and observability.
  • Experience with monitoring tools such as Prometheus, Grafana, or DataDog.

Nice to have

  • Advanced knowledge of Unix/Linux systems, networking fundamentals, and system performance tuning.
  • Familiarity with MySQL, Redis, Nginx, Kafka, Kubernetes, Docker, Hadoop, Spark, Flink, Hive, OLAP, or ClickHouse.

Culture & Benefits

  • Medical, dental, and vision insurance from day one.
  • 401(k) savings plan with company match.
  • Paid parental leave, disability coverage, life insurance, and wellbeing benefits.
  • 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
  • Inclusive workplace with reasonable accommodations available during recruitment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →