Назад
Company hidden
19 часов назад

Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)

136 800 - 259 200$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure): Building and operating large-scale, massively distributed, and fault-tolerant systems with an accent on automation, scalability, monitoring, and incident response. Focus on designing reliable cloud infrastructure, diagnosing production issues, defining SLOs, SLIs, and SLAs, and preventing recurring failures through performance testing and blameless postmortems.

Location: San Jose R&D; fully in-person schedule up to 5 days a week

Salary: $136,800–$259,200 annually

Company

hirify.global Joint Venture operates TikTok-related apps with a focus on data privacy, cybersecurity, national security, and protection of U.S. user data.

What you will do

  • Develop and maintain automation procedures to improve system efficiency and reduce manual intervention.
  • Design, deploy, and operate robust systems in collaboration with software engineering teams.
  • Build for scalability across web traffic, data growth, and large-scale distributed systems.
  • Implement monitoring, metrics, and performance tests to identify bottlenecks and track system health.
  • Participate in on-call rotations, incident management, diagnosis, resolution, and prevention of production issues.
  • Define SLOs, SLIs, and SLAs while conducting sustainable user support and blameless postmortems.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, or a related field and 3+ years of experience.
  • Professional experience as a Site Reliability Engineer, Systems Engineer, or similar software engineering professional.
  • Proficiency in Python, Go, Java, or Shell scripting.
  • Experience with network architecture, database modeling, cloud systems, and large-scale distributed systems.
  • Strong understanding of Linux operating systems and open-source technologies.
  • Availability to work in person in San Jose up to 5 days per week.

Nice to have

  • Experience with Docker, Kubernetes, or equivalent container orchestration platforms.
  • Knowledge of monitoring tools and methodologies such as Prometheus and Grafana.
  • Strong debugging, problem-solving, strategic thinking, and cross-functional communication skills.

Culture & Benefits

  • Medical, dental, and vision insurance from day one.
  • 401(k) savings plan with company match, paid parental leave, disability coverage, and life insurance.
  • Wellbeing benefits, 10 paid holidays, 10 paid sick days, and 17 days of paid personal time.
  • Collaborative, diverse, intellectually curious, and problem-solving-oriented work environment.
  • On-call participation and blameless incident response practices.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →