Назад
Company hidden
6 дней назад

Site Reliability Engineer (Cloud Infrastructure)

122 574 - 259 200$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer (Cloud Infrastructure): Building and operating reliable cloud, networking, and physical infrastructure platforms for TikTok's US region with an accent on automation, monitoring, distributed systems, and compliance. Focus on designing large-scale infrastructure tooling, managing production incidents and capacity, and ensuring service availability through SRE practices and disaster recovery.

Location: San Jose, United States; fully in-person schedule up to 5 days a week

Salary: $122,574–$259,200 annually, plus potential discretionary bonuses, incentives, and restricted stock units.

Company

A technology organization focused on protecting U.S. user data, applications, algorithms, and the content ecosystem through data privacy and cybersecurity programs.

What you will do

  • Provision physical servers and maintain the U.S. physical network and infrastructure.
  • Design, develop, and maintain automation and monitoring tools for large-scale infrastructure.
  • Collaborate with engineering teams to design, deploy, operate, and improve scalable services.
  • Monitor system health, conduct performance testing, and manage production incidents and service-level objectives.
  • Perform on-call operations, change management, capacity planning, disaster recovery, documentation, and process improvements.
  • Collaborate with vendors and international colleagues on hardware, networks, platforms, assurance, and compliance.

Requirements

  • Proficiency in one or more programming languages, such as Python, Go, Java, or C++.
  • Strong knowledge of Linux operating systems and open-source technologies.
  • Experience with network architecture and troubleshooting, database modeling, cloud systems, and large-scale distributed systems.
  • Knowledge of monitoring tools and methodologies, including Prometheus, Grafana, AIOPS, APM, and disaster recovery.
  • Experience designing and building automation and tools for large-scale systems.
  • Experience building solutions with AWS, GCP, Azure, or other cloud services.

Nice to have

  • Expertise with Kubernetes, Elasticsearch, ClickHouse, message queues, OpenTSDB, service mesh, MySQL, Redis, or similar technologies.
  • Master's degree in Computer Science, Engineering, or a related field.

Culture & Benefits

  • On-site work is expected up to 5 days per week.
  • Medical, dental, and vision insurance from the first day.
  • 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
  • Wellbeing benefits, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
  • Inclusive workplace focused on creativity, collaboration, continuous learning, and operational impact.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →