Назад
Company hidden
8 дней назад

Sr GPU Cloud K8S Expert (SRE SME)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr GPU Cloud K8S Expert (SRE SME) (GPU Cloud/Kubernetes): Designing, deploying, and operating an AI-operated GPU cloud control plane for topology-aware scheduling, self-healing, and multi-tenant workloads with an accent on GPU cluster operations, Bare-Metal-as-a-Service, and automated remediation. Focus on building Kubernetes operators and CRDs, automating GPU fault recovery, and meeting cluster availability and job-completion SLOs at 100–10,000 GPU scale.

Location: Remote within San Jose, CA or Austin, TX

Company

hirify.global provides Bitcoin mining infrastructure, AI computational infrastructure, data center operations, and cloud capabilities for artificial intelligence workloads.

What you will do

  • Design, deploy, and operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPU scale.
  • Manage Nvidia GPU Operator, device plugins, MIG configuration, GPU time-slicing, topology-aware scheduling, NVLink locality, and network rail affinity.
  • Build CRDs and integrate AI frameworks including Slurm on Kubernetes, Ray on Kubernetes, and Kubeflow.
  • Implement multi-tenant isolation and Bare-Metal-as-a-Service provisioning, onboarding, lifecycle management, and reclamation.
  • Develop Terraform infrastructure modules, SLI/SLO monitoring, incident automation, and operational workflows using Prometheus, Grafana, Alertmanager, and PagerDuty.
  • Automate GPU node failure detection, draining, cordoning, tainting, workload rescheduling, and AIOps-driven remediation.

Requirements

  • 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
  • Deep knowledge of Nvidia GPU Operator, device plugins, GPU scheduling, topology-aware scheduling, and GPU-specific resource management.
  • Experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
  • Hands-on experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
  • Proficiency in Terraform, Helm, GitOps workflows, ArgoCD or Flux, and programming in Go or Python for operator and CRD development.
  • Strong SRE background covering SLI/SLO frameworks, incident management, capacity planning, monitoring at scale, AIOps remediation, and runbook-as-code.

Culture & Benefits

  • Full-time role focused on operating an AI-managed GPU cloud control plane.
  • Work includes turning human SRE interventions into autonomous platform workflows.
  • Success is measured by automated GPU fault recovery without customer impact, external tenant self-service, and published cluster and job-completion SLOs.
  • hirify.global operates data centers across the United States, Norway, Bhutan, and Ethiopia, with headquarters in Singapore.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →