Назад
9 дней назад

Site Reliability Engineer, Lead

298 000 - 368 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Site Reliability Engineer, Lead (C++/SRE): Leading reliability engineering for planet-scale software release pipelines supporting autonomous vehicles, with an accent on observability, incident management, capacity planning, and production operations. Focus on investigating novel failures, improving distributed-system robustness, and designing automation and architecture for reliable software delivery.

Location: Mountain View, California, United States; hybrid work schedule

Salary: $298,000–$368,000 USD per year, plus eligibility for an annual bonus, equity incentive plan, and company benefits.

Company

Waymo develops autonomous driving technology and operates fully autonomous ride-hail services powered by the Waymo Driver.

What you will do

  • Lead the reliability strategy and execution for critical software release pipelines.
  • Drive incident response, retrospectives, on-call rotations, disaster recovery, and business continuity planning.
  • Harden production software, develop operational playbooks, and improve on-call metrics and post-mitigation workstreams.
  • Lead engineering analysis and response to novel, rare, or unexpected events during the development and release cycle.
  • Promote reliability engineering practices across the Onboard, Eval, and Simulator organizations.
  • Contribute to system software architecture focused on robustness and debuggability.

Requirements

  • Master’s degree or PhD in Computer Science, Engineering, or a related technical field.
  • 10+ years of experience with large-scale production software systems, especially machine-learned systems.
  • 10+ years of hands-on coding experience with C++.
  • Deep expertise in large-scale distributed systems, observability, monitoring, and incident management.
  • Strong background in cloud infrastructure, production operations, automation, and modern DevOps/SRE tooling.
  • Experience organizing on-call rotations, influencing without authority, and improving reliability through data-driven problem-solving.

Nice to have

  • Experience leading hybrid hardware/software production systems.
  • Experience leading teams working on deep machine learning and large-scale models.
  • Experience driving change in engineering organizations of 1,000+ people.

Culture & Benefits

  • Work on software delivery systems supporting a growing fleet of autonomous vehicles.
  • Collaborate across the Onboard, Eval, and Simulator organizations.
  • Eligibility for a discretionary annual bonus program.
  • Eligibility for an equity incentive plan and company benefits program.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →