8 дней назад
Sr GPU Cloud K8S Expert (SRE SME)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr GPU Cloud K8S Expert (SRE SME) (GPU Cloud/Kubernetes): Designing, deploying, and operating an AI-operated GPU cloud control plane for topology-aware scheduling, self-healing, and multi-tenant workloads with an accent on GPU cluster operations, Bare-Metal-as-a-Service, and automated remediation. Focus on building Kubernetes operators and CRDs, automating GPU fault recovery, and meeting cluster availability and job-completion SLOs at 100–10,000 GPU scale.
Location: Remote within San Jose, CA or Austin, TX
Company
provides Bitcoin mining infrastructure, AI computational infrastructure, data center operations, and cloud capabilities for artificial intelligence workloads.
What you will do
- Design, deploy, and operate production Kubernetes clusters optimized for GPU workloads at 100–10,000 GPU scale.
- Manage Nvidia GPU Operator, device plugins, MIG configuration, GPU time-slicing, topology-aware scheduling, NVLink locality, and network rail affinity.
- Build CRDs and integrate AI frameworks including Slurm on Kubernetes, Ray on Kubernetes, and Kubeflow.
- Implement multi-tenant isolation and Bare-Metal-as-a-Service provisioning, onboarding, lifecycle management, and reclamation.
- Develop Terraform infrastructure modules, SLI/SLO monitoring, incident automation, and operational workflows using Prometheus, Grafana, Alertmanager, and PagerDuty.
- Automate GPU node failure detection, draining, cordoning, tainting, workload rescheduling, and AIOps-driven remediation.
Requirements
- 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
- Deep knowledge of Nvidia GPU Operator, device plugins, GPU scheduling, topology-aware scheduling, and GPU-specific resource management.
- Experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
- Hands-on experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
- Proficiency in Terraform, Helm, GitOps workflows, ArgoCD or Flux, and programming in Go or Python for operator and CRD development.
- Strong SRE background covering SLI/SLO frameworks, incident management, capacity planning, monitoring at scale, AIOps remediation, and runbook-as-code.
Culture & Benefits
- Full-time role focused on operating an AI-managed GPU cloud control plane.
- Work includes turning human SRE interventions into autonomous platform workflows.
- Success is measured by automated GPU fault recovery without customer impact, external tenant self-service, and published cluster and job-completion SLOs.
- operates data centers across the United States, Norway, Bhutan, and Ethiopia, with headquarters in Singapore.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →