Назад
Company hidden
5 дней назад

Staff Site Reliability Engineer (Autonomous Vehicles)

172 000 - 300 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Site Reliability Engineer (Autonomous Vehicles) (SRE, cloud infrastructure, Kubernetes): Building a centralized reliability engineering practice and production automation for systems used to build, validate, release, and operate autonomous-vehicle software with an accent on SLOs, error budgets, observability, and incident response. Focus on establishing reliability standards, service catalogs, policy-as-code, failure-mode analysis, and reusable automation across distributed infrastructure.

Location: Sunnyvale, California or Austin, Texas; employees living within a 50-mile radius of an office must report onsite at least three times per week. The role may be eligible for relocation benefits.

Salary: $172,000–$300,000 per year, plus bonus potential and benefits.

Company

hirify.global develops automotive technologies and autonomous-driving systems guided by a vision of zero crashes, zero emissions, and zero congestion.

What you will do

  • Co-found and operationalize a centralized SRE practice, including adoption paths, governance, and measurable reliability outcomes.
  • Build shared tooling for service catalogs, SLOs, error budgets, readiness evidence, and toil reduction.
  • Define risk-based production reliability architecture, minimum standards, golden paths, and exception processes.
  • Establish incident severity, command, response, and blameless post-incident practices with tracked corrective actions.
  • Partner with service owners to classify critical services, map dependencies, identify failure modes, and define meaningful SLIs and SLOs.
  • Turn repeated manual operational work into reusable software while protecting engineering capacity.

Requirements

  • Significant experience as an SRE, Production Engineer, or Infrastructure Software Engineer operating large-scale distributed systems.
  • Experience taking an SRE or operational excellence program from 0 to 1 or improving reliability in an organization with uneven maturity.
  • Strong software engineering skills in at least one general-purpose language, such as Go, Python, Java, or C++.
  • Deep knowledge across Linux, networking, Kubernetes, cloud infrastructure, CI/CD, or observability.
  • Hands-on experience with SLIs, SLOs, error budgets, actionable alerts, and on-call health.
  • Ability to build consensus and drive standards adoption without taking ownership away from service teams.

Nice to have

  • Experience with autonomous vehicles, robotics, safety-relevant systems, automotive software, or other high-consequence production environments.
  • Experience with hybrid cloud and on-premises environments, data or ML platforms, service catalogs, reliability scorecards, or policy-as-code.
  • Experience facilitating game days, failure injection, regional failover, or disaster-recovery exercises.

Culture & Benefits

  • Health, dental, and vision insurance options.
  • Health Savings Account and Flexible Spending Accounts.
  • Retirement savings plan, life insurance, and sickness and accident benefits.
  • Paid vacation and holidays, tuition assistance, and an employee assistance program.
  • GM vehicle discounts and performance-based incentive pay.

Hiring process

  • Applicants may be required to complete role-related assessments and pre-employment screening where applicable.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →