1 день назад
Senior Site Reliability Engineer (AI/LLM)
187 040 - 359 720$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Site Reliability Engineer (AI/LLM): Operating and scaling realtime and batch data services and pipelines with an accent on observability, incident response, and AI-powered automation. Focus on designing LLM-powered operational tools, leading high-severity incidents, and building reliable distributed systems with Kubernetes and Kafka.
Location: San Jose, United States; fully in-person schedule up to 5 days a week
Salary: $187,040–$359,720 annually, plus potential discretionary bonuses, incentives, and restricted stock units.
Company
A USDS joint venture focused on data privacy, cybersecurity, national security, and trust and safety for applications serving American users.
What you will do
- Manage day-to-day operations of realtime and batch data services and pipelines, including SLA/SLO/SLI management, deployments, performance tuning, and troubleshooting.
- Design and deploy AI agents and LLM-powered automation for incident response, root cause analysis, monitoring, and runbook matching.
- Build internal administration platforms and automation that reduce operational toil and improve engineering velocity.
- Participate in on-call rotations providing 24-hour coverage, lead incident response and postmortems, and serve as Incident Commander for high-severity issues.
- Own service lifecycle activities including design, capacity planning, launch reviews, deployment, operation, and refinement.
- Partner with Product and Development teams to embed observability, scalability, and disaster recovery into new features, while mentoring junior engineers.
Requirements
- Bachelor’s degree or higher in Computer Science or a related technical discipline.
- At least 5 years of professional experience in SRE, DevOps, or infrastructure engineering.
- Experience integrating AI and LLM APIs into internal workflows for log analysis, alert contextualization, or diagnostic assistance.
- Deep knowledge of Unix/Linux internals, networking fundamentals, and distributed systems architecture.
- Experience designing and scaling observability stacks with Prometheus, Grafana, or Datadog.
- Ability to troubleshoot complex production issues across the entire technology stack.
Nice to have
- Experience building agentic workflows and orchestration frameworks for incident triage and runbook matching.
- Advanced use of AI-assisted development tools for infrastructure-as-code and automated documentation.
- Strong technical proficiency with Kubernetes and big data technologies such as Kafka.
Culture & Benefits
- Fully in-person work supports rapid alignment, real-time decision-making, and integrated execution.
- Medical, dental, and vision insurance from day one, plus a 401(k) plan with company match.
- Paid parental leave, short- and long-term disability coverage, life insurance, and wellbeing benefits.
- 10 paid holidays, 10 paid sick days, and 17 days of paid personal time with increasing accrual by tenure.
- Inclusive workplace committed to diversity, reasonable accommodations, curiosity, and continuous improvement.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 часов назад
Site Reliability Engineer, Platform Responsibility - USDS (AI/LLM)
129 960 - 246 240$
19 часов назад
Site Reliability Engineer, Tech Infra - USDS (Cloud Infrastructure)
136 800 - 259 200$
1 день назад
Staff Site Reliability Engineer (GCP/Kubernetes)
112 500 - 187 500$
5 дней назад
Senior Software Engineer — Observability & IRM (Kubernetes)
124 900 - 250 000$
2 дня назад
Systems Engineer (AI Infrastructure)
140 000 - 225 000$
3 дня назад
Sr. DevOps Engineer II (AI)
119 000 - 221 000$