Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Software Engineer, Chaos Engineering (Kubernetes): Building zonal-resilience automation, fault-injection systems, and distributed reliability tooling for Datadog’s production environment with an accent on Kubernetes, controlled failure handling, and safe recovery. Focus on designing reversible production experiments, orchestrating gamedays, connecting findings to verified remediation, and applying AI to identify and close resilience gaps.
Location: Paris, France; hybrid workplace
Company
Datadog provides an observability and security platform that unifies visibility across applications, infrastructure, data, models, and security.
What you will do
- Build zonal-resilience automation for safe workload evacuations, switchovers, and recovery.
- Design fault-injection systems for infrastructure and applications, including incident replay and controlled resilience experiments.
- Implement blast-radius controls, kill switches, validation mechanisms, and rollback paths for production experiments.
- Develop agents and automation that propose failure scenarios, triage experiment results, and connect findings to remediation verification.
- Lead gamedays from scenario design and execution through documented findings, remediation tracking, and validation.
- Design distributed systems such as gRPC services, Kubernetes controllers, and shared platform components while mentoring engineers.
Requirements
- Strong fundamentals in distributed systems, including consistency, failure modes, backpressure, idempotency, quorum, retries, and recovery.
- Understanding of Kubernetes workload lifecycles, including pods, controllers, scheduling, draining, and eviction.
- Experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
- Ability to communicate technical decisions through design documents, runbooks, postmortems, and technical discussions.
- Ability to collaborate across engineering teams, understand unfamiliar systems, identify failure modes, and drive resilience improvements.
Nice to have
- Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems.
Culture & Benefits
- Develop expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure.
- Work on reliability systems operating across Datadog’s production environment.
- Explore automated approaches to fault injection, zonal resilience, incident reproduction, and AI-assisted reliability workflows.
- Collaborate with infrastructure, database, observability, and service engineering teams.
- Mentor engineers and contribute to technical designs, engineering practices, and platform strategy.
- Benefits vary based on the country of employment and the nature of employment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →