2 дня назад
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineering Team Lead (Principal SRE, Automotive AI) (Kubernetes, Azure, Observability): Leading the reliability function for cloud-native automotive AI platforms with an accent on SLI/SLO governance, incident escalation, observability, and production automation. Focus on setting technical direction, designing sustainable on-call operations, improving high-availability systems, and embedding reliability into the software delivery lifecycle.
Location: Remote - USA
Company
develops automotive AI solutions, including cloud-native voice, gesture, and gaze systems used by global automakers.
What you will do
- Lead the Site Reliability Engineering function and set technical direction across a distributed team.
- Own the reliability roadmap, SLI/SLO/SLA standards, on-call model, and high-risk production change approvals.
- Act as the Tier 2 technical escalation point for major incidents and drive actionable blameless postmortems.
- Lead Production Readiness and NFR reviews while embedding reliability practices into the SDLC.
- Set the direction for observability, metrics, dashboards, alerting, and CI/CD automation.
- Partner with development, DevOps, platform, and operations teams on architecture and shared infrastructure.
Requirements
- 8+ years of hands-on experience in SRE, DevOps, or cloud platform roles, including team leadership or functional ownership.
- Hands-on experience with Kubernetes, Docker, Istio, and public cloud platforms, primarily Azure as well as AWS and Google Cloud.
- Experience with observability tools such as Zabbix, Prometheus, and Grafana.
- Experience with CI/CD pipelines and infrastructure as code, including Terraform or Flux.
- Proficiency in Python, Go, Shell, or another scripting or programming language.
- Strong UNIX/Linux, networking, system configuration, performance debugging, and English communication skills.
Nice to have
- Experience leading distributed or multi-site technical teams.
- Background in high-availability service design, redundancy, failover, and blast-radius reduction.
- Experience with Loki, Thanos, Jira, or Confluence.
- Experience in automotive, embedded, or latency-sensitive production environments.
Culture & Benefits
- Remote work is available for this position.
- Annual bonus opportunity and equity awards for eligible positions and levels.
- Medical, dental, vision, life, and disability insurance coverage.
- Paid time off, paid holidays, and company contribution to an RRSP.
- Blameless incident management and a focus on sustainable on-call operations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
8 дней назад
Engineering Manager, SRE (AI)
5 дней назад
Site Reliability Engineer, Tech Lead (AI)
5 дней назад
Senior DevOps / Site Reliability Engineer (SRE) (Cybersecurity)
165 000 - 215 000$
8 дней назад
Senior Staff Site Reliability Engineer
232 338 - 290 422$
3 дня назад
Site Reliability Engineer III
9 дней назад
Staff Site Reliability Engineer (AI)
252 000 - 308 000$