9 дней назад
Sr. GPU Cloud K8S Expert (SRE SME)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr. GPU Cloud K8S Expert (SRE SME) (GPU cloud/Kubernetes/SRE): Operating and scaling production Kubernetes clusters for GPU workloads from 100 to 10,000 GPUs with an accent on GPU scheduling, multi-tenant isolation, bare-metal provisioning, and automated remediation. Focus on designing CRDs and control-plane workflows, building BMaaS self-service infrastructure, and meeting cluster availability and job-completion SLOs without customer impact.
Location: Singapore, SG / Penang, MY
Company
is a technology company providing Bitcoin mining solutions, AI cloud capabilities, ASIC hardware, and computing infrastructure.
What you will do
- Own production Kubernetes clusters optimized for GPU workloads at a scale of 100–10,000 GPUs.
- Configure NVIDIA GPU Operator, device plugins, MIG, GPU time-slicing, topology-aware scheduling, and GPU resource management.
- Develop CRDs and integrate AI frameworks including Slurm on Kubernetes, Ray on Kubernetes, and Kubeflow.
- Build multi-tenant isolation, bare-metal provisioning, tenant onboarding, lifecycle automation, and reclamation workflows.
- Implement Terraform infrastructure modules, monitoring, SLIs/SLOs, incident management, and automated GPU node failure handling.
- Make the Kubernetes control plane safe for automated remediation and convert SRE runbooks into executable workflows.
Requirements
- 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
- Deep knowledge of NVIDIA GPU Operator, device plugins, GPU scheduling, topology-aware scheduling, and GPU-specific resource management.
- Experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
- Hands-on experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom tooling.
- Proficiency in Terraform, Helm, GitOps workflows, ArgoCD or Flux, Prometheus, Grafana, and large-scale alerting.
- Strong Go or Python programming skills for operator and CRD development, plus an SRE background in SLOs, incidents, and capacity planning.
Nice to have
- Experience integrating autoscaler or remediator loops with Kubernetes.
- Strong AIOps aptitude and a runbook-as-code mindset.
Culture & Benefits
- Inclusive environment valuing authenticity and diverse perspectives.
- Start-up spirit within a fast-growing technology company.
- Autonomy, personal accountability, and opportunities to contribute to new systems and projects.
- Training, mentoring, welfare benefits, and professional development opportunities.
- Opportunity to contribute to digital asset and AI cloud infrastructure.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →