Назад
Company hidden
2 месяца назад

Lead Site Reliability Engineer, Incident Management

Формат работы
remote (только India)/hybrid
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
India
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Lead Site Reliability Engineer, Incident Management (SRE/Cloud): Leading production reliability and major incident response across a global SaaS platform with an accent on automation, observability, scalability, and cross-functional operational leadership. Focus on designing self-healing platforms, reducing MTTD and MTTR, improving SLOs and error budgets, and restoring complex production outages.

Location: Pune; hybrid or remote work environment

Company

Global SaaS platform supporting production reliability, scalability, and customer experience.

What you will do

  • Lead enterprise-wide critical incident response as Incident Commander.
  • Own restoration strategies for complex production outages, including executive communications and customer-impact assessments.
  • Lead root cause analysis and systemic reliability improvements.
  • Design self-healing platforms and automation while improving observability, SLOs, SLIs, and error budgets.
  • Drive capacity planning, operational readiness, architecture reviews, and cross-functional operational standards.
  • Mentor Senior and Lead SREs and shape the SRE strategy and reliability roadmap.

Requirements

  • Bachelor’s degree or equivalent.
  • 8–12+ years of experience in SRE or production operations.
  • Strong Linux, Kubernetes, cloud, networking, and database skills.
  • Expertise in Python, Go, Bash, or Java.
  • Experience leading major incidents and mentoring engineers.
  • Excellent communication and leadership skills; availability for 24x7 production support and on-call leadership.

Nice to have

  • Large-scale SaaS experience.
  • Multi-cloud expertise.
  • Chaos Engineering experience.
  • Cloud or Kubernetes certifications.
  • ITIL certification.

Culture & Benefits

  • Full-time role with hybrid or remote work.
  • Cross-functional collaboration with Engineering, Infrastructure, DevOps, Cloud Operations, Database Engineering, Security, Product Engineering, Customer Support, and Executive Leadership.
  • Opportunity to drive automation adoption, reliability culture, and measurable improvements in availability and resiliency.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →