10 часов назад
Site Reliability Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (AI): Building reliability tools, automation, observability and remediation workflows for autonomous vehicle software and fleet operations with an accent on distributed systems, production reliability and safe vehicle operations. Focus on debugging incidents end to end, defining SLIs and SLOs, improving monitoring and alerting, and driving lasting resilience across cloud, networking and vehicle-adjacent environments.
Location: Leonberg, Germany; hybrid working model with in-person collaboration in the office, vehicle workshops and labs
Company
is building an AI platform for autonomous driving that enables vehicles to learn from real-world experience and adapt across complex urban environments.
What you will do
- Monitor live system health using logs, metrics, traces, alerts and reliability dashboards.
- Investigate production incidents end to end, from hypothesis formation and data analysis through mitigation and root-cause resolution.
- Participate in on-call rotations and post-incident reviews.
- Build tools and automation that reduce manual operational work and accelerate issue detection and resolution.
- Develop service-level indicators and objectives, improve monitoring and alerting, and create automated checks and remediation workflows.
- Influence system design and improve reliability across software, infrastructure, networking and fleet operations.
Requirements
- Hands-on experience in Site Reliability Engineering or a similar role supporting live production systems.
- Strong programming skills in Python, C++ or Rust.
- Solid Linux fundamentals and experience with cloud platforms, containers, Kubernetes and CI/CD.
- Knowledge of observability practices and tools such as Datadog, Prometheus, Grafana, OpenTelemetry or Splunk.
- Understanding of SLIs, SLOs and error budgets.
- Ability to debug complex distributed systems across software, networking, operating systems and hardware-adjacent environments.
Culture & Benefits
- Hybrid work combining office collaboration with focused remote work.
- Relocation support and visa sponsorship where applicable.
- Market-benchmarked salaries, meaningful equity and location-dependent benefits.
- Learning and development budgets for training, conferences and professional growth.
- Health insurance, dental coverage, enhanced parental leave, retirement or pension benefits where applicable, therapy access, wellbeing partnerships and team socials.
- Ownership and influence in a company whose processes and ways of working are still being developed.
Hiring process
- Initial recruiter call.
- Competency interviews followed by deep-dive technical interviews.
- Final interview, with interview formats explained in advance and scheduling adapted to availability.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →