13 дней назад
Platform Engineer – Cloud & Observability (GCP)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Platform Engineer – Cloud & Observability (GCP): Building and operating scalable enterprise platform services on GCP and Kubernetes with an accent on observability, infrastructure automation, and developer self-service. Focus on defining SLIs and SLOs, troubleshooting complex production issues, and improving reliability through Terraform, ArgoCD, GitOps, and automated operational practices.
Location: Bengaluru, India — office-based
Company
develops API-first digital banking, account opening, and branch solutions for financial institutions across digital, remote, and in-person channels.
What you will do
- Design, build, and operate scalable, secure, and reliable platform services on GCP and Kubernetes/GKE.
- Define observability standards covering metrics, logs, traces, dashboards, alerts, Golden Signals, SLIs, and SLOs.
- Develop reusable platform capabilities, self-service solutions, operational standards, and runbooks for engineering teams.
- Implement Infrastructure as Code with Terraform and automation using Python, Shell, or similar technologies.
- Support application delivery through ArgoCD, GitOps, and CI/CD practices.
- Troubleshoot production issues, participate in incident response and root cause analysis, and improve reliability, scalability, and operational efficiency.
Requirements
- 6+ years of experience in Platform Engineering, Cloud Engineering, SRE, DevOps, Observability Engineering, or a related discipline.
- Strong hands-on experience with GCP, cloud architecture, Kubernetes/GKE, Terraform, ArgoCD, GitOps, and CI/CD.
- Knowledge of cloud networking, IAM, compute, storage, load balancing, and cloud-native platforms.
- Experience with observability concepts and production solutions such as Dynatrace, Datadog, Splunk, AppDynamics, Prometheus, or Grafana.
- Experience with dashboards, alerts, monitoring standards, distributed tracing, automation, incident management, and root cause analysis.
- Strong communication, collaboration, holistic troubleshooting, and problem-solving skills.
Nice to have
- Experience building reusable platform and self-service capabilities for engineering teams.
- Experience with OpenTelemetry, Prometheus, Grafana, or similar technologies.
- Experience integrating observability platforms with cloud platforms, CI/CD pipelines, or incident management workflows.
- Experience with cloud migration, modernization, highly available architectures, or regulated financial services environments.
Culture & Benefits
- Platform ownership and a strong reliability mindset.
- Focus on automation, standardization, self-service, developer experience, and operational simplicity.
- Close collaboration with SRE, Software Engineering, Security, and Architecture teams.
- Emphasis on proactive identification of reliability and performance risks across the cloud and application ecosystem.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →