2 дня назад
SRE (Infrastructure) Lead Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
SRE (Infrastructure) Lead Engineer (Azure/Kubernetes): Operating and evolving the cloud and Kubernetes platform powering retail robotics, store operations, telemetry, and customer-facing services with an accent on reliability engineering, infrastructure as code, observability, security, and cost control. Focus on leading incident response, automating deployments and operations, strengthening cloud-to-edge resilience, and guiding a hands-on Infra/SRE team.
Location: Tokyo, Japan; hybrid workplace
Company
develops retail robotics products, including robots, smart shelves, store operations systems, and supporting SaaS services.
What you will do
- Own reliability practices for production cloud and Kubernetes systems, including SLOs, SLIs, error budgets, alerting, and operational readiness.
- Lead infrastructure incident response, act as Incident Commander when needed, and drive blameless postmortems and corrective actions.
- Design, operate, and improve AKS and Kubernetes environments, including autoscaling, ingress, TLS, RBAC, cluster upgrades, and workload scheduling.
- Maintain Infrastructure as Code with Azure Bicep, Terraform, and Terragrunt, and evolve GitOps deployments with Argo CD, Kustomize, and Helm.
- Improve CI/CD, observability, security controls, disaster recovery, and cloud cost management across development, staging, infrastructure, and production.
- Lead and mentor 3–6 Infra/SRE/Platform engineers while collaborating with Backend, Frontend, Robotics, Security, Product, and Operations teams.
Requirements
- 5+ years of professional infrastructure, platform, SRE, DevOps, backend infrastructure, or cloud engineering experience.
- 2+ years in a senior, lead, or technical leadership role responsible for production reliability or infrastructure direction.
- Hands-on production experience with a major cloud platform such as Azure, AWS, or GCP, plus production Kubernetes experience.
- Infrastructure as Code experience with Terraform, Bicep, CloudFormation, Pulumi, CDK, or similar tools.
- Experience designing and operating production CI/CD pipelines, observability systems, and incident response practices.
- Strong Linux, networking, DNS, TLS, cloud identity, scripting, communication, and cross-functional collaboration skills.
Nice to have
- Azure production experience with AKS, Azure Monitor, Application Insights, Log Analytics, Key Vault, Entra ID, managed identities, or Azure RBAC.
- Experience with Argo CD, Kustomize, Helm, cert-manager, Traefik, ingress-nginx, Grafana, Loki, Alloy, or Prometheus-style monitoring.
- Experience with GitHub Actions OIDC, workload identity, policy-as-code, drift detection, cloud guardrails, or compliance infrastructure.
- Background in robotics, IoT, retail operations, edge computing, telemetry systems, or other cyber-physical production environments.
- Japanese language skills.
Culture & Benefits
- Hands-on technical leadership with direct operation of production systems.
- Opportunity to build a stronger SRE and platform engineering practice for retail robotics.
- Focus on pragmatic reliability, security, automation, standardization, and financially sustainable cloud architecture.
- Opportunity to shape infrastructure practices across cloud resources, Kubernetes, CI/CD, observability, and incident response.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 часа назад
Senior Site Reliability Engineer (AWS)
5 дней назад
DevOps Tech Lead (Cloud)
6 дней назад
IT Staff Systems Engineer - VM (Infrastructure Automation & Virtualization)
DeepL
6 дней назад
Senior Platform Engineer (Kubernetes)
DeepL
5 дней назад
Developer Experience Engineer (AI)
4 дня назад