7 дней назад
Site Reliability Engineer (Cloud Infrastructure)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer (Cloud Infrastructure): Leading incident response and building reliable, scalable, and highly available infrastructure for warehousing IT systems with an accent on monitoring, automation, and root cause analysis. Focus on optimizing system architecture, implementing preventive measures, and resolving complex performance and availability issues across cloud and containerized environments.
Location: Manila Net Park Office, Manila, Philippines. Standard five-day workweek with rotating weekend and holiday on-call assignments. Work-from-home arrangements may be available.
Company
Procter & Gamble produces globally recognized consumer brands and operates in approximately 70 countries.
What you will do
- Lead incident response efforts and resolve critical incidents while minimizing downtime and user impact.
- Establish incident management processes, including communication, coordination, documentation, and post-incident reviews.
- Conduct root cause analysis and implement preventive measures to improve system resilience.
- Design, optimize, and automate highly available, scalable, and fault-tolerant infrastructure.
- Implement and manage monitoring, alerting, and observability solutions for real-time system visibility.
- Collaborate with software engineers, DevOps specialists, subject matter experts, customers, and users.
Requirements
- 3–6 years of experience in system administration, including Linux/Unix, cloud platforms such as AWS, Azure, or GCP, and SAP.
- Experience with configuration management and infrastructure-as-code frameworks such as Terraform.
- Proficiency in at least one programming language, such as Python or C#, and scripting for automation.
- Knowledge of networking, load balancing, DNS, databases, SQL, containerization, and orchestration technologies such as Docker and Kubernetes.
- Experience with monitoring and observability tools such as Prometheus and Grafana, incident response, root cause analysis, and secure systems implementation.
- Strong troubleshooting, communication, collaboration, prioritization, and customer-support skills.
Nice to have
- Experience with Warehousing Management Systems such as RTCIS or PrIME, or with warehousing operations.
Culture & Benefits
- Flexible work schedule with available work-from-home arrangements.
- Performance bonus through the STAR program.
- Health insurance, fitness support, and an Employee Assistance Program.
- Opportunities for technical growth, leadership, mentoring, and collaboration on transformative projects.
- Culture of continuous learning, knowledge sharing, and innovation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Site Reliability Engineer (FinTech)
23 часа назад
Platform Site Reliability Engineer (SRE)
5 дней назад
Lead, Reliability & Service Health Engineering
5 дней назад
Senior Enterprise Observability Operations Engineer
6 дней назад
Observability Engineer (Fintech)
7 дней назад