4 дня назад
Site Reliability Engineering, Engineer II (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineering, Engineer II (Kubernetes): Operating and improving the reliability, scalability, and performance of production services across Kubernetes-based cloud and hybrid infrastructure with an accent on automation, observability, and operational excellence. Focus on troubleshooting distributed systems, supporting CI/CD and GitOps workflows, building infrastructure-as-code tooling, and responding to production incidents in a rotating on-call schedule.
Location: Palermo, Buenos Aires, Argentina; hybrid, 3 days per week onsite
Company
develops Experience Cloud, a SaaS platform for managing customer, employee, patient, candidate, and resident experiences.
What you will do
- Operate and support production services running in Kubernetes-based cloud and hybrid environments.
- Improve application reliability, scalability, performance, and operational maturity in collaboration with software engineering teams.
- Build automation, reusable tooling, self-service solutions, and infrastructure-as-code configurations to reduce manual operational work.
- Support CI/CD and GitOps deployment workflows and monitor system health through observability and alerting platforms.
- Troubleshoot infrastructure and application issues, participate in incident response and root cause analysis, and improve operational standards.
- Use AI-assisted engineering tools and automation platforms to improve troubleshooting, productivity, and service reliability.
Requirements
- At least 2 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Operations, or related roles.
- Experience supporting production environments on Kubernetes or other containerized platforms and cloud infrastructure such as AWS, OCI, or GCP.
- Linux systems administration and troubleshooting experience, plus scripting or programming with Python, Bash, or Go.
- Familiarity with CI/CD pipelines, Git-based workflows, networking fundamentals, and distributed-systems troubleshooting.
- Fluent English, both oral and written.
- Ability to work onsite three days per week in the Buenos Aires vicinity and participate in a rotating production on-call schedule.
Nice to have
- Experience with GitOps, ArgoCD, Terraform, Prometheus, Grafana, Loki, or OpenTelemetry.
- Experience operating hybrid-cloud or multi-region services and using rolling, canary, or blue/green release strategies.
- Knowledge of incident management, operational best practices, security, and compliance in production environments.
- Experience with AI-assisted development, automation, or operational tooling.
Culture & Benefits
- Work on a global SaaS platform supporting production services used by customers.
- Collaborate closely with software engineering teams on reliability and operational improvements.
- Participate in a fast-paced environment focused on automation, continuous improvement, and operational efficiency.
- Equal opportunity employment and reasonable accommodations for applicants with disabilities.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →