3 дня назад
Senior/Staff Platform Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior/Staff Platform Engineer (Kubernetes): Building, operating, and evolving large-scale production infrastructure with an accent on Kubernetes platforms, cloud systems, automation, and reliability engineering. Focus on diagnosing multi-layer infrastructure failures, improving observability and resilience, and leading complex platform initiatives from design through production.
Location: Remote for candidates based in Brazil, Mexico, or Canada; Pacific Hours, 8:00 AM–5:00 PM PST, with on-call every 4–5 weeks.
Company
provides technology and engineering services involving complex cloud, platform, and infrastructure environments.
What you will do
- Design, build, operate, and improve production Kubernetes platforms, including cluster architecture, networking, security, scaling, and reliability.
- Diagnose complex failures across Kubernetes, containers, Linux, networking, cloud infrastructure, and distributed systems.
- Develop production tooling, internal services, APIs, and automation using Go, Python, or Java.
- Own reliability engineering, incident response, SLOs, SLIs, observability, disaster recovery, and operational improvements.
- Build infrastructure as code with Terraform and improve CI/CD, deployment, migration, rollback, and production validation workflows.
- Lead ambiguous infrastructure initiatives, communicate technical decisions to stakeholders, and mentor engineers.
Requirements
- 10+ years of experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates.
- Hands-on experience building and operating production Kubernetes platforms, not only deploying applications to existing clusters.
- Production programming experience in Go, Python, or Java.
- Strong Linux, networking, cloud, Terraform, automation, troubleshooting, incident response, and distributed-systems experience.
- Experience with observability, CI/CD infrastructure, Docker, high availability, capacity planning, disaster recovery, and production resilience.
- Must currently be based in Brazil, Mexico, or Canada and be available during Pacific working hours.
Nice to have
- Cloud or infrastructure migration experience, including cutover and rollback strategies.
- Large-scale or multi-cluster Kubernetes, hybrid cloud, on-premises, virtualized, or bare-metal experience.
- Kubernetes controllers or operators, advanced networking, service mesh, mTLS, workload identity, or multi-cloud experience.
- Disaster recovery testing, performance engineering, platform tooling, infrastructure security, or compliance experience.
- Technical leadership, mentoring, or customer-facing consulting experience.
Culture & Benefits
- Fully remote work within the stated candidate locations.
- Autonomous work with ownership of complex infrastructure initiatives.
- Direct collaboration with customer and internal engineering teams.
- Distributed, highly technical environment with a focus on clear technical communication and ownership.
- On-call participation approximately every 4–5 weeks.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →