Назад
22 дня назад

Staff Site Reliability Engineer (Automotive)

251 000 - 310 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (C++/Java/Python): Building and operating reliable infrastructure for autonomous vehicle fleet services, including depot logistics, vehicle state systems, observability, and deployment automation, with an accent on fault tolerance, availability, and performance. Focus on leading incident response, resolving complex distributed-system failures, managing technical debt, and setting architecture standards across engineering organizations.

Location: On site in Mountain View or San Francisco, California, United States. Once- or twice-yearly travel to the Bay Area is preferred but not required.

Salary: $251,000–$310,000 USD base salary per year, plus eligibility for an annual bonus, equity incentive plan, and company benefits.

Company

Waymo develops autonomous driving technology and operates a fully autonomous ride-hailing service powered by the Waymo Driver.

What you will do

  • Build and operate reliable systems for autonomous vehicle operations, including depot logistics, automation flows, and critical vehicle-state infrastructure.
  • Own end-to-end availability and performance for core fleet services and maintain sufficient usable vehicle supply for demand.
  • Design and implement architecture, telemetry, deployment, observability, and automation improvements for mission-critical fleet services.
  • Lead troubleshooting of complex reliability issues, technical debt management, incident response, and sustainable on-call operations.
  • Set architectural guidelines and lead cross-functional initiatives that integrate projects into reusable components and services.
  • Mentor engineers, develop leadership, and promote operational excellence through blameless retrospectives.

Requirements

  • 8+ years of experience architecting and maintaining mission-critical systems in C++, Java, or Python.
  • Deep expertise in software reliability and resolving complex, high-impact system failures.
  • Experience leading the reliability program of a complex system such as autonomous systems, cloud infrastructure, or a large-scale web service.
  • Ability to develop and communicate strategy across organizations, influence key decisions, and lead Engineering–Development initiatives.
  • Experience building team leadership through delegation, recruiting, and mentoring engineers.
  • Bachelor’s degree in a relevant field, or 10+ years of comparable experience in a high-growth environment with leadership experience.

Nice to have

  • Experience leading multiple large projects or a mission-critical program from concept to delivery.
  • Experience translating business needs into site reliability initiatives and managing technical debt and system evolution.
  • Experience leading architectural investigations and resolving failures in distributed systems.
  • Master’s or PhD in Computer Science, Computer Engineering, or a related field.

Culture & Benefits

  • Participation in a sustainable on-call rotation with a focus on continuous improvement.
  • Blameless retrospectives and operational excellence practices.
  • Discretionary annual bonus program and equity incentive plan, subject to eligibility.
  • Company benefits program, subject to eligibility.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →