Назад
Company hidden
5 дней назад

K8 Site Reliability SME (Kubernetes)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
K8 Site Reliability SME (Kubernetes) (GPU Kubernetes/AIOps): Designing, deploying, and operating production Kubernetes control planes for AI GPU cloud workloads with an accent on topology-aware scheduling, automated remediation, and multi-tenant isolation. Focus on building GPU fault drain and rescheduling workflows, automating bare-metal tenant lifecycle management, and meeting cluster availability and job-completion SLOs.

Location: Remote within San Jose, California or Austin, Texas

Company

hirify.global provides Bitcoin mining solutions and AI computational infrastructure, including data center, equipment management, operations, and cloud capabilities.

What you will do

  • Design, deploy, and operate production Kubernetes clusters for GPU workloads ranging from 100 to 10,000 GPUs.
  • Configure the Nvidia GPU Operator, device plugin, MIG, GPU time-slicing, and topology-aware scheduling across NVLink domains and network rails.
  • Develop CRDs and integrate Kubernetes with Slurm, Ray, and Kubeflow for AI workload lifecycle management.
  • Build multi-tenant isolation using namespaces, network policies, quotas, RBAC, and pod security standards.
  • Automate BMaaS provisioning, tenant onboarding, lifecycle management, reclamation, and GPU node failure recovery.
  • Operate monitoring and reliability systems with Prometheus, Grafana, Alertmanager, PagerDuty, SLIs/SLOs, runbook automation, and incident reviews.

Requirements

  • 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
  • Deep knowledge of the Nvidia GPU Operator, device plugin, GPU scheduling, topology-aware scheduling, and GPU resource management.
  • Hands-on experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
  • Experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
  • Proficiency in Terraform, Helm, GitOps workflows, ArgoCD or Flux, and programming in Go or Python for operator and CRD development.
  • Strong SRE experience with SLI/SLO frameworks, incident management, capacity planning, monitoring at scale, and automated remediation.

Culture & Benefits

  • Work on an AI-operated GPU cloud and infrastructure supporting AI computational workloads.
  • Build an AIOps execution surface where remediation workflows can act safely without manual pager intervention.
  • Apply a runbook-as-code approach so operational playbooks become executable platform workflows.
  • Equal employment opportunities are provided in accordance with applicable country, state, and local laws.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →