обновлено 3 месяца назад
Lead Site Reliability Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer (SLS): Building and leading SRE for multi-tenant distributed storage services that enhance MongoDB Atlas with an accent on reliability, durability, and operational safety. Focus on defining SLOs, shaping capacity plans, executing multi-year roadmap, and guiding architectural design reviews.
Location: Dublin, Ireland; hybrid working model
Company
provides a globally distributed, multi-cloud data platform through Atlas for customers modernizing workloads and building AI-enabled applications.
What you will do
- Build and lead a team of 6–8 SREs, supporting career growth, performance management, team culture, and blocker removal.
- Define the technical vision and roadmap for performant, multi-tenant distributed storage systems.
- Establish SLOs, shape capacity plans, and ensure reliability, durability, and operational safety for the Atlas storage layer.
- Lead architectural design reviews, review pull requests, and guide the team through complex operational challenges.
- Collaborate with engineering leaders as the primary liaison for the Storage Layer Services SRE team.
Requirements
- 10+ years of experience working with software and operating distributed systems, including 2+ years managing engineering teams.
- Deep familiarity with Kubernetes ecosystems, containerization, and infrastructure-as-code tooling such as Terraform, Crossplane, or Operators.
- Experience operating or supporting stateful storage or database systems at scale, including durability, consistency, and recovery trade-offs.
- Ability to translate complex business and engineering requirements into actionable, phased technical roadmaps.
- Strong technical communication, ownership, accountability, empathy, and a customer-focused mindset.
- Must be based in Dublin for the hybrid working model.
Nice to have
- Experience leading migrations from legacy storage stacks to new multi-tenant storage architectures.
- Experience planning and executing large-scale data and workload migrations with strict availability and durability requirements.
- Experience managing infrastructure across AWS, Google Cloud, or Microsoft Azure.
- Experience designing secure, multi-tenant runtime environments at scale.
Culture & Benefits
- Supportive and enriching environment focused on employee growth and business impact.
- Employee affinity groups and programs supporting employee wellbeing.
- Fertility assistance and generous parental leave.
- Accommodation support is available during the application and interview process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →