Назад
Company hidden
9 дней назад

Sr. GPU Cloud K8S Expert (SRE SME)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. GPU Cloud K8S Expert (SRE SME) (GPU cloud/Kubernetes/SRE): Operating and scaling production Kubernetes clusters for GPU workloads from 100 to 10,000 GPUs with an accent on GPU scheduling, multi-tenant isolation, bare-metal provisioning, and automated remediation. Focus on designing CRDs and control-plane workflows, building BMaaS self-service infrastructure, and meeting cluster availability and job-completion SLOs without customer impact.

Location: Singapore, SG / Penang, MY

Company

hirify.global is a technology company providing Bitcoin mining solutions, AI cloud capabilities, ASIC hardware, and computing infrastructure.

What you will do

  • Own production Kubernetes clusters optimized for GPU workloads at a scale of 100–10,000 GPUs.
  • Configure NVIDIA GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, and GPU resource management.
  • Develop CRDs and integrate AI frameworks including Slurm on Kubernetes, Ray on Kubernetes, and Kubeflow.
  • Build multi-tenant isolation, bare-metal provisioning, tenant onboarding, lifecycle automation, and reclamation workflows.
  • Implement Terraform infrastructure modules, monitoring, SLIs/SLOs, incident management, and automated GPU node failure handling.
  • Make the Kubernetes control plane safe for automated remediation and convert SRE runbooks into executable workflows.

Requirements

  • 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
  • Deep knowledge of NVIDIA GPU Operator, device plugins, GPU scheduling, topology-aware scheduling, and GPU-specific resource management.
  • Experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
  • Hands-on experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom tooling.
  • Proficiency in Terraform, Helm, GitOps workflows, ArgoCD or Flux, Prometheus, Grafana, and large-scale alerting.
  • Strong Go or Python programming skills for operator and CRD development, plus an SRE background in SLOs, incidents, and capacity planning.

Nice to have

  • Experience integrating autoscaler or remediator loops with Kubernetes.
  • Strong AIOps aptitude and a runbook-as-code mindset.

Culture & Benefits

  • Inclusive environment valuing authenticity and diverse perspectives.
  • Start-up spirit within a fast-growing technology company.
  • Autonomy, personal accountability, and opportunities to contribute to new systems and projects.
  • Training, mentoring, welfare benefits, and professional development opportunities.
  • Opportunity to contribute to digital asset and AI cloud infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →