Назад
Company hidden
7 дней назад

Senior Site Reliability Engineer (SRE)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior/lead
Английский
b2
Страна
Philippines
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Site Reliability Engineer (SRE): Leading incident response and building reliable, scalable, and highly monitored systems for warehousing IT operations with an accent on automation, observability, and service reliability. Focus on managing SLOs and SLIs, optimizing architecture for fault tolerance, resolving complex incidents, and mentoring the SRE team.

Location: Manila Net Park Office, Philippines; flexible work-from-home arrangements are available, with a standard five-day workweek and rotating weekend and holiday on-call coverage.

Company

Procter & Gamble produces and operates globally recognized consumer brands across approximately 70 countries.

What you will do

  • Lead the SRE Incident Response team and mentor engineers across technical excellence, knowledge sharing, and professional development.
  • Coordinate rapid response, resolution, documentation, and root cause analysis for critical incidents.
  • Improve system availability, scalability, performance, and fault tolerance through resilient architecture, automation, monitoring, and alerting.
  • Manage SLOs, SLIs, schedules, reporting, and operational coverage across the week.
  • Collaborate with software engineers, DevOps specialists, subject matter experts, customers, and users to address reliability needs.
  • Report incident, performance, and reliability insights to the IT Operations Director.

Requirements

  • At least 7 years of experience in software engineering, software development, SRE, DevOps, or technical consulting.
  • Knowledge of Linux/Unix system administration, cloud platforms such as AWS, Azure, or GCP, and SAP.
  • Experience with configuration management, infrastructure as code, Terraform, scripting, and at least one programming language such as Python or C#.
  • Understanding of networking, load balancing, DNS, containerization, Kubernetes, databases, and SQL.
  • Experience with monitoring and observability tools such as Prometheus and Grafana, incident response, root cause analysis, and secure systems.
  • Strong troubleshooting, communication, collaboration, leadership, scheduling, and high-pressure incident management skills.

Nice to have

  • Experience with warehousing management systems such as RTCIS or PrIME, or with warehousing operations.

Culture & Benefits

  • Flexible work schedule with available work-from-home arrangements.
  • Performance bonus through the STAR program.
  • Health insurance, fitness support, and an Employee Assistance Program.
  • Opportunities to lead transformative technology projects and develop technical and leadership expertise.
  • Predictable scheduling is supported while maintaining customer coverage and service reliability.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →