16 часов назад
Staff Platform Site Reliability Engineer (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Platform Site Reliability Engineer (Kubernetes): Building Index Cloud, a multi-tenant globally distributed compute platform, with an accent on Kubernetes infrastructure, infrastructure-as-code, platform APIs, and large-scale distributed systems. Focus on designing multi-datacenter systems, deploying safely across thousands of servers, driving platform architecture, and improving developer self-service.
Location: London, hybrid; office attendance is expected on Tuesdays, Wednesdays and Thursdays
Company
operates a global advertising supply-side platform that processes more than 700 billion real-time auctions daily through proprietary, privacy-first ad technology.
What you will do
- Design and deliver multi-tenant Kubernetes infrastructure across bare-metal servers and public cloud.
- Build infrastructure-as-code frameworks that deploy changes across thousands of servers.
- Develop standard libraries, internal SDKs, and platform APIs for engineering teams.
- Solve distributed systems challenges involving sub-millisecond real-time bidding, multi-datacenter consistency, fleet deployment, and large-scale load balancing.
- Drive technical direction through RFCs, design reviews, tooling standards, security decisions, and system architecture.
- Mentor engineers and collaborate with Cloud Platform Operations, SRE, Network, Security, and Software Engineering teams.
Requirements
- 8+ years of experience in platform engineering, SRE, infrastructure engineering, or DevOps.
- Deep knowledge of Linux internals, including kernel tuning, networking, observability, and security.
- Strong Kubernetes expertise covering cluster lifecycle, networking, storage, RBAC, and multi-cluster environments across bare metal and cloud.
- Experience with Terraform, Ansible, and GitOps tools such as ArgoCD.
- Proficiency in Go or Python for building libraries, SDKs, and platform APIs.
- Strong L2-L7 networking fundamentals, load balancing, DNS, service discovery, and cross-team technical strategy.
Nice to have
- Experience with Ceph, Hadoop, Spark, HBase, Kafka, Prometheus, Grafana, ELK, Mimir, Loki, or Tempo.
- Knowledge of Vault, certificate management, and access control at scale.
- Experience with hybrid cloud architectures combining AWS or GCP with on-premises infrastructure.
- Experience managing bare-metal infrastructure in globally distributed data centers.
Culture & Benefits
- Health, dental, and vision plans for employees and dependents.
- Paid time off, health days, personal obligation days, and flexible work schedules.
- Retirement matching, equity packages, parental leave, and commuter benefits where available.
- Well-being allowance, fitness discounts, wellness activities, and an employee assistance program.
- Volunteer time off, donation matching, continuous learning resources, town halls, and community-led events.
- Inclusive workplace with support for accessibility accommodations.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →