1 день назад
Site Reliability Engineer II
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Site Reliability Engineer II (Kubernetes/cloud infrastructure): Operating and improving the reliability, scalability, and performance of production services across Kubernetes-based cloud and hybrid infrastructure with an accent on automation, observability, and operational excellence. Focus on troubleshooting distributed systems, building infrastructure-as-code and CI/CD tooling, and supporting incident response through a rotating on-call schedule.
Location: Cuauhtémoc, Mexico; candidates in the Mexico City vicinity are prioritized. Hybrid schedule with 3 days per week onsite.
Company
Experience Management SaaS platform serving organizations across customer, employee, patient, and resident experiences.
What you will do
- Operate and support production services running in Kubernetes-based cloud and hybrid environments.
- Improve application reliability, scalability, performance, and operational maturity in collaboration with software engineering teams.
- Troubleshoot infrastructure, application, and distributed-system issues and participate in incident response and root cause analysis.
- Build automation, self-service tooling, and reusable operational solutions to reduce manual work and operational toil.
- Support CI/CD and GitOps deployment workflows and maintain infrastructure-as-code configurations.
- Monitor system health, availability, and performance using observability and alerting platforms.
Requirements
- 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Operations, or a related role.
- Experience supporting production environments on Kubernetes or other containerized platforms and cloud infrastructure such as AWS, OCI, or GCP.
- Linux administration and troubleshooting experience.
- Scripting or programming experience with Python, Bash, or Go.
- Familiarity with CI/CD pipelines, Git-based workflows, networking fundamentals, and distributed-system troubleshooting.
- Ability to work hybrid with 3 days per week onsite and participate in a rotating on-call schedule. Professional working proficiency in written and spoken English is required.
Nice to have
- Experience with GitOps, ArgoCD, Terraform, Prometheus, Grafana, Loki, or OpenTelemetry.
- Experience operating hybrid-cloud or multi-region services and using rolling, canary, or blue/green deployment strategies.
- Experience with incident management, production security and compliance, or AI-assisted engineering and operational tooling.
- Strong communication, collaboration, automation, and process-improvement skills.
Culture & Benefits
- Work on a global SaaS platform supporting production services used by customers worldwide.
- Engineering practices emphasize automation, platform thinking, self-service, and reducing operational toil.
- AI-assisted engineering workflows are used to improve productivity and service reliability.
- Equal opportunity workplace with accommodations available during the application process.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →