Назад
Company hidden
5 дней назад

Senior AI Infrastructure Engineer, Kubernetes

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Infrastructure Engineer, Kubernetes (AI infrastructure/Kubernetes): Building and operating production-grade Kubernetes platforms for GPU-accelerated bare-metal AI infrastructure with an accent on cluster lifecycle, networking, storage, security, and observability. Focus on designing multi-tenant clusters, integrating NVIDIA GPU and high-performance networking technologies, and solving complex distributed-systems reliability challenges.

Location: San Francisco Bay Area, United States

Company

hirify.global Technologies develops and operates energy-efficient AI infrastructure and the hirify.global AI Cloud GPU platform for training and deploying AI models.

What you will do

  • Define the Kubernetes reference architecture for management and workload clusters, including lifecycle, multi-tenancy, workload isolation, and failure-domain design.
  • Build backend services, APIs, controllers, operators, and automation for provisioning, upgrading, scaling, and retiring Kubernetes clusters.
  • Engineer bare-metal Kubernetes deployment workflows and operate networking, storage, security, observability, and disaster-recovery capabilities.
  • Integrate NVIDIA GPU and Network Operators, scheduling, telemetry, quotas, and topology-aware placement for accelerated AI workloads.
  • Establish GitOps, CI/CD, progressive delivery, rollback, policy, and software-supply-chain controls.
  • Set engineering standards, provide technical sign-off, mentor engineers, and lead resolution of complex platform failures across infrastructure teams.

Requirements

  • 7+ years of infrastructure, systems, or platform engineering experience, including substantial ownership of production Kubernetes platforms.
  • At least 3 years at senior staff, principal, or equivalent level, with experience designing and operating highly available, large-scale, multi-cluster Kubernetes platforms.
  • Deep knowledge of Kubernetes internals, Linux systems, networking, storage, security, governance, and cluster performance.
  • Strong software engineering skills in Go and/or Rust, with practical Python and Bash experience building operators, controllers, webhooks, CLIs, or platform services.
  • Experience with GPU-enabled Kubernetes, NVIDIA GPU Operator, accelerator scheduling, RDMA networking, distributed storage, observability, and disaster recovery.
  • CKA-level expertise is expected; a bachelor's degree or equivalent practical engineering experience is required.

Nice to have

  • CKA, CKS, or relevant cloud-native certifications.
  • Experience with Cluster API, kubeadm, Redfish, PXE, Ironic, Metal3, Multus, SR-IOV, BGP, InfiniBand, or RoCE.

Culture & Benefits

  • Full-time employment with reporting to the Head of AI Platform.
  • Work on sustainable AI infrastructure and energy-efficient GPU cloud technology.
  • Inclusive workplace encouraging applications from candidates of all backgrounds.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →