5 дней назад
Senior AI Infrastructure Engineer, Kubernetes
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior AI Infrastructure Engineer, Kubernetes (AI infrastructure/Kubernetes): Building and operating production-grade Kubernetes platforms for GPU-accelerated bare-metal AI infrastructure with an accent on cluster lifecycle, networking, storage, security, and observability. Focus on designing multi-tenant clusters, integrating NVIDIA GPU and high-performance networking technologies, and solving complex distributed-systems reliability challenges.
Location: San Francisco Bay Area, United States
Company
Technologies develops and operates energy-efficient AI infrastructure and the AI Cloud GPU platform for training and deploying AI models.
What you will do
- Define the Kubernetes reference architecture for management and workload clusters, including lifecycle, multi-tenancy, workload isolation, and failure-domain design.
- Build backend services, APIs, controllers, operators, and automation for provisioning, upgrading, scaling, and retiring Kubernetes clusters.
- Engineer bare-metal Kubernetes deployment workflows and operate networking, storage, security, observability, and disaster-recovery capabilities.
- Integrate NVIDIA GPU and Network Operators, scheduling, telemetry, quotas, and topology-aware placement for accelerated AI workloads.
- Establish GitOps, CI/CD, progressive delivery, rollback, policy, and software-supply-chain controls.
- Set engineering standards, provide technical sign-off, mentor engineers, and lead resolution of complex platform failures across infrastructure teams.
Requirements
- 7+ years of infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms.
- At least 3 years at senior staff, principal, or equivalent level, with experience designing and operating highly available, large-scale, multi-cluster Kubernetes platforms.
- Deep knowledge of Kubernetes internals, Linux systems, networking, storage, security, governance, and cluster performance.
- Strong software engineering skills in Go and/or Rust, with practical Python and Bash experience building operators, controllers, webhooks, CLIs, or platform services.
- Experience with GPU-enabled Kubernetes, NVIDIA GPU Operator, accelerator scheduling, RDMA networking, distributed storage, observability, and disaster recovery.
- CKA-level expertise is expected; a bachelor's degree or equivalent practical engineering experience is required.
Nice to have
- CKA, CKS, or relevant cloud-native certifications.
- Experience with Cluster API, kubeadm, Redfish, PXE, Ironic, Metal3, Multus, SR-IOV, BGP, InfiniBand, or RoCE.
Culture & Benefits
- Full-time employment with reporting to the Head of AI Platform.
- Work on sustainable AI infrastructure and energy-efficient GPU cloud technology.
- Inclusive workplace encouraging applications from candidates of all backgrounds.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
Principal Member of Technical Staff (AI Infrastructure)
200 000 - 350 000$
9 дней назад
AI Platform Engineer (AI)
163 900 - 215 000$
11 дней назад
Lead Engineer, Platform Engineering & Reliability (AI)
6 дней назад
Senior Platform Engineer (Kubernetes)
9 дней назад
Engineering Manager (AI)
210 000 - 300 000$
10 дней назад
Senior Software Engineer (Infrastructure)
170 000 - 220 000$