Назад
Company hidden
1 день назад

Site Reliability Engineer, AI Infrastructure (AI)

129 960 - 246 240$
Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer, AI Infrastructure (AI): Building and operating globally distributed, fault-tolerant infrastructure for recommendation and search systems with an accent on production ownership, observability, automation, and service reliability. Focus on designing distributed systems, leading infrastructure migrations, managing capacity and SLAs/SLOs, and resolving systemic causes of incidents in a highly secure environment.

Location: Seattle, United States; fully in-person schedule up to 5 days a week

Salary: $129,960–$246,240 annually, with potential additional bonuses, incentives, and restricted stock units.

Company

A technology joint venture focused on data privacy, cybersecurity, national security, and secure U.S. user data and applications.

What you will do

  • Partner with engineering and product teams across system design, architecture reviews, deployment, operations, and continuous service refinement.
  • Build tools, platforms, and automation that improve reliability, scalability, R&D efficiency, and operational workflows.
  • Monitor service health, latency, and key metrics for large-scale, multi-region systems.
  • Lead infrastructure migrations and architecture upgrades in a highly restricted compliance environment.
  • Manage capacity, resource allocation, stability optimization, error attribution, and SLA/SLO compliance.
  • Drive incident management, blameless postmortems, and systemic root-cause resolution.

Requirements

  • Bachelor’s degree or equivalent practical experience in computer science, software engineering, or a related technical field.
  • At least 1 year of hands-on SRE, DevOps, or systems engineering experience with large-scale, highly reliable systems.
  • Strong knowledge of Linux, system performance, networking fundamentals, and Bash or shell scripting.
  • Programming experience in at least one of Go, Python, C/C++, or Java.
  • Experience designing, troubleshooting, and maintaining complex distributed systems in dynamic production environments.
  • Familiarity with CI/CD practices and automated deployment pipelines.

Nice to have

  • Experience with Kubernetes, service mesh architectures, cloud platforms, and Infrastructure-as-Code such as Terraform.
  • Experience with Prometheus, Grafana, distributed tracing, or other observability tools.
  • Experience building self-service tools and automation, or applying LLMs and agentic AI to operational workflows.

Culture & Benefits

  • On-site collaboration in a highly secure and strictly isolated infrastructure environment.
  • Medical, dental, and vision insurance from day one, plus a 401(k) plan with company match.
  • Paid parental leave, disability coverage, life insurance, and wellbeing benefits.
  • 10 paid holidays, 10 paid sick days, and 17 days of paid personal time, with increasing accruals by tenure.
  • Commitment to inclusive recruitment and reasonable accommodations.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →