Senior Platform Engineer, Cloud Infrastructure (Go)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior Platform Engineer, Cloud Infrastructure (Go): Build and operate the cloud-native Kubernetes platform that products run on, with an accent on production reliability, Kubernetes internals, and service mesh capabilities. Focus on designing multi-region platform architecture, writing Go services/controllers, and driving SLOs, incident response, and observability to keep dependent engineering teams running safely.
Location: Remote (Pacific Hours: 8:00 AM – 5:00 PM PST)
Company
builds cloud-native products and relies on a Kubernetes platform for production operations.
What you will do
- Design, build, and operate production Kubernetes clusters, including networking, workload isolation, and multi-region topologies.
- Work with Kubernetes internals (resource quotas, scheduling behavior, NetworkPolicy) and build custom controllers/operators.
- Develop and maintain Go services, controllers, and middleware that extend the platform, including HTTP/REST and gRPC interfaces.
- Own reliability: lead incident response, perform root-cause analysis, write postmortems, and implement durable fixes.
- Define and drive SLOs with actionable alerting; troubleshoot using logs, metrics, traces, and profiling.
- Build platform delivery workflows with Infrastructure as Code, CI/CD, and GitOps; maintain reusable IaC modules.
Requirements
- 6+ years of professional experience in software/platform/infrastructure/SRE, including significant time operating production distributed systems.
- Proven experience building and operating production Kubernetes platforms (not just deploying to Kubernetes).
- Production experience writing Go as the primary day-to-day language.
- Strong Kubernetes internals knowledge: networking (CNI), NetworkPolicy, resource management, and cluster behavior under load.
- Hands-on service mesh experience (Istio/Envoy/Linkerd or similar) including mTLS and workload identity.
- Experience with observability tooling (Prometheus, Grafana, OpenTelemetry, PromQL) and production debugging/performance profiling.
Nice to have
- Experience with Go testing frameworks (Ginkgo, Gomega).
- Experience building Kubernetes controllers/operators and API-server extensions.
- Additional strength in Python.
- Experience with identity and access management (SSO, Keycloak, OIDC, SAML) and/or secret management (Vault or cloud equivalent).
- Experience with compliance frameworks (SOC 2, GDPR) and/or designing for high availability and disaster recovery across regions.
Culture & Benefits
- Fully remote role with coverage aligned to Pacific Hours (8:00 AM – 5:00 PM PST).
- Hands-on platform engineering where value is measured by platform capability and reliability.
- Ownership of ambiguous, multi-quarter initiatives through production delivery.
- Cross-functional collaboration with product, security, and infrastructure teams, plus mentoring and engineering standards.
Hiring process
- Interviews focused on Kubernetes/platform engineering, Go services/controllers, reliability ownership, and observability/SLO practices.
- Discussion of experience taking ambiguous initiatives to production and collaborating across teams.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →