5 дней назад
Team Lead, Site Reliability Engineering (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Team Lead, Site Reliability Engineering (SRE): Leading reliability engineering for a highly distributed financial platform with an accent on high-level reliability architecture, observability, automation, and incident management. Focus on designing self-healing systems, scaling Kubernetes infrastructure, defining SLOs, and mentoring SRE engineers through complex production challenges.
Location: Lagos, Nigeria
Company
is an African financial platform providing payments, banking, credit, cross-border, and business management tools for businesses and individuals.
What you will do
- Set the technical direction for the SRE team and design self-healing systems.
- Define reliability standards, production readiness reviews, observability strategies, and automation practices.
- Establish end-to-end system visibility through logging, tracing, metrics, monitoring, and actionable alerting.
- Lead, mentor, and develop Senior and Associate SREs through code reviews and technical workshops.
- Act as the escalation point for major incidents and improve incident management and root-cause analysis processes.
- Partner with Engineering Managers and Product Leads to define service level objectives aligned with business goals.
Requirements
- At least 6 years of experience in SRE or Backend Engineering, including 2 years in a Lead or Senior/Staff role mentoring others.
- Expert-level proficiency in Java, Go, Rust, or Python.
- Strong knowledge of distributed systems patterns, scalable architectures, and microservices troubleshooting.
- Deep expertise with Google Cloud Platform or AWS, including running Kubernetes at scale and troubleshooting infrastructure issues.
- Proven experience defining observability strategies and architecting telemetry stacks from instrumentation to alerting.
- Strong communication skills and the ability to de-escalate high-pressure incident response situations.
Culture & Benefits
- People-first culture focused on well-being, inclusion, respect, and open communication.
- Learning and development environment with knowledge sharing, training, and internal technical talks.
- Salary, pension, health insurance, annual bonus, and additional benefits.
Hiring process
- Preliminary phone call with a recruiter.
- Technical interview with the Hiring Manager.
- Behavioural and technical interview with an Executive team member.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
Senior Site Reliability Engineer (Kubernetes)
5 дней назад
Senior Site Reliability Engineer (SRE)
5 дней назад
Staff Site Reliability Engineer (AI)
5 дней назад
Senior Site Reliability Engineer (SRE)
Deimos
7 дней назад
Senior Site Reliability Engineer
2 дня назад