1 час назад
Associate Site Reliability Engineer (SRE) - Multicloud Platform (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Associate Site Reliability Engineer (SRE) - Multicloud Platform (Kubernetes): Supporting and scaling a reliable cloud platform across distributed AWS and GCP environments with an accent on Kubernetes infrastructure, automation, and observability. Focus on diagnosing complex reliability issues, implementing SLO-gated deployment and runbook automation, and maintaining continuous coverage through follow-the-sun incident response.
Location: Auckland, New Zealand; hybrid work with at least 50% of each quarter spent in the office or in the field
Company
is a Fortune 500 company providing an AI-powered platform for managing people, money, and agents.
What you will do
- Diagnose and resolve reliability and performance issues across distributed cloud environments and Kubernetes clusters.
- Design automation that reduces operational toil and improves platform efficiency at scale.
- Develop SLIs, support SLO achievement, and extend observability, runbook automation, and deployment capabilities.
- Partner with platform service teams to define SRE standards, benchmarks, and production qualification automation.
- Contribute to incident response, root cause analysis, runbooks, and reliability improvements.
- Participate in a follow-the-sun on-call roster with SRE counterparts across NZT, GMT, and PT.
Requirements
- Some exposure to SRE, DevOps, or cloud-native infrastructure through projects, coursework, certifications, internships, or prior work.
- Basic hands-on experience with AWS, GCP, or Azure and growing familiarity with Kubernetes and CNCF technologies.
- Working knowledge of Linux/Unix fundamentals, Git, GitOps principles, CI/CD, testing, and software development practices.
- Some experience with Go, Python, or Ruby and an interest in writing scripts and automation.
- BSc in Computer Science or a related field, or equivalent practical experience and demonstrable project work.
- Strong communication, documentation, collaboration, and independent problem-solving skills.
Nice to have
- Familiarity with Istio, OPA, Prometheus, Grafana, or ArgoCD.
- Experience with Go and Kubernetes troubleshooting.
Culture & Benefits
- Flexible work combines in-person collaboration with remote flexibility.
- At least 50% of each quarter is spent in the office or with customers, prospects, and partners in the field, depending on the role.
- Global collaboration across cloud platform teams in New Zealand, the United States, and Ireland.
- Autonomous two-week sprint planning and opportunities to contribute to the cloud-native community.
- Reasonable accommodations are available during the application process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →