Назад
обновлено 7 дней назад

Staff Infrastructure Engineer (Kubernetes)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Infrastructure Engineer (Kubernetes): Design and evolve Kubernetes control plane architecture for multi-tenant, multi-region AI compute platform with an accent on scalability, reliability, and operational ownership. Focus on multi-tenant cluster models, regional scaling strategies, networking integration, and production incident resolution.

Location: Remote; must have authorization to work in the United States

Company

TensorWave provides a secure, reliable cloud platform for delivering AI compute at scale.

What you will do

  • Design and evolve Kubernetes control plane architecture across regions and data centers.
  • Define multi-tenant cluster models, isolation boundaries, resource segmentation, and policy enforcement.
  • Own production platform reliability, participate in on-call rotation, and lead incident response and root cause analysis.
  • Manage cluster provisioning, upgrades, scaling, topology, and failure-domain strategies.
  • Design cluster and regional ingress and egress architectures and optimize pod networking, CNI behavior, and high-performance network integration.
  • Improve observability, resilience, and platform alignment with compute, storage, networking, DevOps, and CI/CD infrastructure.

Requirements

  • 7+ years of experience in infrastructure, platform engineering, or distributed systems.
  • Deep experience operating Kubernetes at scale across multiple clusters, regions, or data centers.
  • Strong understanding of Kubernetes internals, including the API server, scheduler, controller manager, and etcd.
  • Strong Linux systems expertise and troubleshooting ability across Kubernetes, container runtimes, and networking.
  • Experience designing control plane architectures and multi-tenant cluster models.
  • Authorization to work in the United States is required.

Nice to have

  • Experience with virtual cluster technologies such as vcluster or Kamaji.
  • Experience supporting GPU workloads, NUMA-aware scheduling, topology-aware workloads, RDMA, or high-throughput networking.
  • Experience with Cilium and observability platforms such as Prometheus and Grafana.
  • Experience in CSP, hyperscale, or equivalent large-scale environments.

Culture & Benefits

  • Fully remote work environment with flexible paid time off and paid holidays.
  • Stock options.
  • 100% paid medical, dental, and vision insurance for employees.
  • Health savings account contributions, flexible spending account, and supplemental insurance options.
  • Short- and long-term disability insurance, life insurance, parental leave, and employee assistance program.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →