7 дней назад
Site Reliability Engineer - Warehousing IT Operations (Cloud Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer - Warehousing IT Operations (Cloud Infrastructure): Leading incident response and maintaining reliable, scalable warehousing IT systems with an accent on monitoring, automation, and fault-tolerant infrastructure. Focus on root cause analysis, optimizing cloud and containerized systems, implementing observability solutions, and coordinating rotating on-call support.
Location: Taguig City, Philippines
Company
Procter & Gamble produces globally recognized consumer brands and operates in approximately 70 countries.
What you will do
- Lead incident response efforts, coordinate communications, resolve critical incidents, and minimize downtime and user impact.
- Conduct root cause analysis, document incidents, implement preventive measures, and improve incident management practices.
- Design, implement, and automate resilient, highly available infrastructure and services for warehousing IT operations.
- Configure monitoring, alerting, and observability solutions to provide real-time system health and performance insights.
- Optimize system architecture and configurations for performance, scalability, fault tolerance, and resource efficiency.
- Collaborate with software engineers, DevOps teams, customers, users, and other stakeholders while mentoring team members.
Requirements
- Knowledge of system administration in Linux/Unix environments, cloud platforms such as AWS, Azure, or GCP, and SAP.
- Experience with configuration management, infrastructure as code, and tools such as Terraform.
- Proficiency in at least one programming language, such as Python or C#, plus scripting for automation.
- Understanding of networking, load balancing, DNS management, containers, Kubernetes, databases, and SQL.
- Familiarity with monitoring and observability tools such as Prometheus and Grafana, along with incident response, root cause analysis, and security best practices.
- Availability for a standard five-day workweek with rotating weekend and holiday on-call assignments.
Nice to have
- Experience with warehousing management systems such as RTCIS or PrIME, or with warehousing operations.
Culture & Benefits
- Work in a cross-functional environment focused on reliability, continuous improvement, and customer support.
- Access opportunities for technical upskilling, knowledge sharing, and mentoring.
- Predictable and manageable scheduling is supported while maintaining incident coverage and system availability.
- Equal opportunity workplace committed to diversity and inclusion.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Site Reliability Engineer (FinTech)
23 часа назад
Platform Site Reliability Engineer (SRE)
5 дней назад
Lead, Reliability & Service Health Engineering
5 дней назад
Senior Enterprise Observability Operations Engineer
6 дней назад
Senior Enterprise Observability Operations Engineer
6 дней назад