2 месяца назад
Lead Site Reliability Engineer, Incident Management
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Lead Site Reliability Engineer, Incident Management (SRE/Cloud): Leading production reliability and major incident response across a global SaaS platform with an accent on automation, observability, scalability, and cross-functional operational leadership. Focus on designing self-healing platforms, reducing MTTD and MTTR, improving SLOs and error budgets, and restoring complex production outages.
Location: Pune; hybrid or remote work environment
Company
Global SaaS platform supporting production reliability, scalability, and customer experience.
What you will do
- Lead enterprise-wide critical incident response as Incident Commander.
- Own restoration strategies for complex production outages, including executive communications and customer-impact assessments.
- Lead root cause analysis and systemic reliability improvements.
- Design self-healing platforms and automation while improving observability, SLOs, SLIs, and error budgets.
- Drive capacity planning, operational readiness, architecture reviews, and cross-functional operational standards.
- Mentor Senior and Lead SREs and shape the SRE strategy and reliability roadmap.
Requirements
- Bachelor’s degree or equivalent.
- 8–12+ years of experience in SRE or production operations.
- Strong Linux, Kubernetes, cloud, networking, and database skills.
- Expertise in Python, Go, Bash, or Java.
- Experience leading major incidents and mentoring engineers.
- Excellent communication and leadership skills; availability for 24x7 production support and on-call leadership.
Nice to have
- Large-scale SaaS experience.
- Multi-cloud expertise.
- Chaos Engineering experience.
- Cloud or Kubernetes certifications.
- ITIL certification.
Culture & Benefits
- Full-time role with hybrid or remote work.
- Cross-functional collaboration with Engineering, Infrastructure, DevOps, Cloud Operations, Database Engineering, Security, Product Engineering, Customer Support, and Executive Leadership.
- Opportunity to drive automation adoption, reliability culture, and measurable improvements in availability and resiliency.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →