Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Infrastructure Engineer (Kubernetes): Design and evolve Kubernetes control plane architecture for multi-tenant, multi-region AI compute platform with an accent on scalability, reliability, and operational ownership. Focus on multi-tenant cluster models, regional scaling strategies, networking integration, and production incident resolution.
Location: Remote; must have authorization to work in the United States
Company
TensorWave provides a secure, reliable cloud platform for delivering AI compute at scale.
What you will do
- Design and evolve Kubernetes control plane architecture across regions and data centers.
- Define multi-tenant cluster models, isolation boundaries, resource segmentation, and policy enforcement.
- Own production platform reliability, participate in on-call rotation, and lead incident response and root cause analysis.
- Manage cluster provisioning, upgrades, scaling, topology, and failure-domain strategies.
- Design cluster and regional ingress and egress architectures and optimize pod networking, CNI behavior, and high-performance network integration.
- Improve observability, resilience, and platform alignment with compute, storage, networking, DevOps, and CI/CD infrastructure.
Requirements
- 7+ years of experience in infrastructure, platform engineering, or distributed systems.
- Deep experience operating Kubernetes at scale across multiple clusters, regions, or data centers.
- Strong understanding of Kubernetes internals, including the API server, scheduler, controller manager, and etcd.
- Strong Linux systems expertise and troubleshooting ability across Kubernetes, container runtimes, and networking.
- Experience designing control plane architectures and multi-tenant cluster models.
- Authorization to work in the United States is required.
Nice to have
- Experience with virtual cluster technologies such as vcluster or Kamaji.
- Experience supporting GPU workloads, NUMA-aware scheduling, topology-aware workloads, RDMA, or high-throughput networking.
- Experience with Cilium and observability platforms such as Prometheus and Grafana.
- Experience in CSP, hyperscale, or equivalent large-scale environments.
Culture & Benefits
- Fully remote work environment with flexible paid time off and paid holidays.
- Stock options.
- 100% paid medical, dental, and vision insurance for employees.
- Health savings account contributions, flexible spending account, and supplemental insurance options.
- Short- and long-term disability insurance, life insurance, parental leave, and employee assistance program.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Staff Software Engineer (Kubernetes)
215 000 - 265 000$
8 дней назад
Senior Cloud Infrastructure Engineer (AWS/Kubernetes)
9 дней назад
Software Engineer (AI Platform)
Mercury
13 дней назад
Software Engineer - Infrastructure (AI)
122 400 - 158 400$
9 дней назад
Senior Cloud Engineer (AWS)
13 дней назад
Platform Engineer (AWS)
80 000 - 100 000€