23 часа назад
Platform Site Reliability Engineer (SRE)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Site Reliability Engineer (SRE) (AWS/DevOps): Building and operating reliable, scalable standard platforms and production applications with an accent on SLOs, observability, automation, and incident response. Focus on capacity planning, complex root-cause analysis, vulnerability and end-of-life governance, and leading technical improvements across Enterprise Platform.
Location: Hybrid role based in Manila, Philippines
Company
provides technology and financial services platforms and supports product engineering teams through its Enterprise Platform group.
What you will do
- Implement SRE practices including error budgets, SLOs, SLIs, monitoring, and alerting.
- Build automation and standard platform capabilities that improve application stability, reliability, scalability, and extensibility.
- Perform capacity planning, system design, troubleshooting, root-cause analysis, incident response, and post-mortem reviews.
- Define and implement reliability and non-functional requirements with software development teams.
- Manage vulnerability and end-of-life governance across products and support deep-dive improvement initiatives for problematic applications.
- Oversee team initiatives, make technical and design decisions, mentor engineers, and contribute to the strategic direction of the SRE function.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Software Engineering, or a related field.
- 5+ years of experience supporting production applications in SRE and/or DevOps roles.
- Experience implementing SLOs, SLIs, observability, and alerting tools such as Datadog or Splunk.
- Knowledge of Windows and/or Linux administration, networking, AWS, middleware, databases, web servers, MQ, and Kafka.
- Experience with automation using Python, Java, shell scripting, Terraform, Chef, Puppet, SQL, or Ansible, plus Docker and Kubernetes.
- Fluent English is essential. Strong problem-solving, collaboration, prioritization, and continuous-improvement skills are required.
Culture & Benefits
- Collaborative and inclusive workplace focused on enabling associates to do their best work.
- Environment that values authenticity, safety, diverse perspectives, and employee contribution.
- Human review is included in employment decisions when AI-based hiring tools are used.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
6 дней назад
Senior Site Reliability Engineer (FinTech)
7 дней назад
Senior Site Reliability Engineer (SRE)
7 дней назад
Site Reliability Engineer (Cloud Infrastructure)
7 дней назад
Site Reliability Engineer - Warehousing IT Operations (Cloud Infrastructure)
5 дней назад
Lead, Reliability & Service Health Engineering
7 дней назад