Назад
Company hidden
1 день назад

Senior Site Reliability Engineer (AI/LLM)

187 040 - 359 720$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (AI/LLM): Operating and scaling realtime and batch data services and pipelines with an accent on observability, incident response, and AI-powered automation. Focus on designing LLM-powered operational tools, leading high-severity incidents, and building reliable distributed systems with Kubernetes and Kafka.

Location: San Jose, United States; fully in-person schedule up to 5 days a week

Salary: $187,040–$359,720 annually, plus potential discretionary bonuses, incentives, and restricted stock units.

Company

A USDS joint venture focused on data privacy, cybersecurity, national security, and trust and safety for applications serving American users.

What you will do

  • Manage day-to-day operations of realtime and batch data services and pipelines, including SLA/SLO/SLI management, deployments, performance tuning, and troubleshooting.
  • Design and deploy AI agents and LLM-powered automation for incident response, root cause analysis, monitoring, and runbook matching.
  • Build internal administration platforms and automation that reduce operational toil and improve engineering velocity.
  • Participate in on-call rotations providing 24-hour coverage, lead incident response and postmortems, and serve as Incident Commander for high-severity issues.
  • Own service lifecycle activities including design, capacity planning, launch reviews, deployment, operation, and refinement.
  • Partner with Product and Development teams to embed observability, scalability, and disaster recovery into new features, while mentoring junior engineers.

Requirements

  • Bachelor’s degree or higher in Computer Science or a related technical discipline.
  • At least 5 years of professional experience in SRE, DevOps, or infrastructure engineering.
  • Experience integrating AI and LLM APIs into internal workflows for log analysis, alert contextualization, or diagnostic assistance.
  • Deep knowledge of Unix/Linux internals, networking fundamentals, and distributed systems architecture.
  • Experience designing and scaling observability stacks with Prometheus, Grafana, or Datadog.
  • Ability to troubleshoot complex production issues across the entire technology stack.

Nice to have

  • Experience building agentic workflows and orchestration frameworks for incident triage and runbook matching.
  • Advanced use of AI-assisted development tools for infrastructure-as-code and automated documentation.
  • Strong technical proficiency with Kubernetes and big data technologies such as Kafka.

Culture & Benefits

  • Fully in-person work supports rapid alignment, real-time decision-making, and integrated execution.
  • Medical, dental, and vision insurance from day one, plus a 401(k) plan with company match.
  • Paid parental leave, short- and long-term disability coverage, life insurance, and wellbeing benefits.
  • 10 paid holidays, 10 paid sick days, and 17 days of paid personal time with increasing accrual by tenure.
  • Inclusive workplace committed to diversity, reasonable accommodations, curiosity, and continuous improvement.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →