Назад
Company hidden
7 часов назад

Tech Lead, AI Infra Site Reliability Engineer (AI)

187 040 - 438 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Tech Lead, AI Infra Site Reliability Engineer (AI) (Recommendation and Search infrastructure): Building and operating large-scale, globally distributed, observable, fault-tolerant systems for recommendation and search engines with an accent on reliability, scalability, automation, and production ownership. Focus on designing distributed systems, optimizing cloud resources and cluster SLAs, improving observability, and managing incident response and postmortems.

Location: San Jose, United States; fully in-person schedule up to 5 days a week

Salary: $187,040–$438,000 annually, with potential discretionary bonuses, incentives, and restricted stock units.

Company

hirify.global operates TikTok-related apps and focuses on protecting U.S. user data, national security, cybersecurity, and the U.S. content ecosystem.

What you will do

  • Improve the full lifecycle of recommendation and search systems, from system design consulting and launch reviews through deployment, operation, and refinement.
  • Develop tools and software that improve service reliability, scalability, operational automation, and R&D efficiency.
  • Build availability for large-scale services deployed across global data centers.
  • Plan, manage, and optimize cloud resource utilization while maintaining SLAs for large-scale clusters.
  • Measure and monitor availability, latency, and overall service health.
  • Lead sustainable incident response and postmortem practices.

Requirements

  • Bachelor’s degree or higher in Computer Science or a related field, plus 5+ years of relevant experience.
  • Experience with SRE for large-scale systems requiring high reliability and scalability.
  • Linux and networking operations experience.
  • Programming experience in at least one of Python, Perl, Go, C, or C++.
  • Experience designing, analyzing, and troubleshooting large-scale distributed systems.
  • Familiarity with CI/CD environments, effective communication, and strong ownership.

Nice to have

  • SRE experience with recommendation systems, search platforms, or advertising platforms.
  • Familiarity with the machine learning lifecycle and inference engines.

Culture & Benefits

  • Medical, dental, and vision insurance from day one.
  • 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
  • 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with accrual increasing by tenure.
  • Work in an in-person environment focused on speed, alignment, agility, curiosity, humility, resilience, and continuous iteration.
  • Reasonable accommodations are available during recruitment for candidates with disabilities, pregnancy, sincerely held religious beliefs, or other legally protected reasons.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →