5 дней назад
K8 Site Reliability SME (Kubernetes)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
K8 Site Reliability SME (Kubernetes) (GPU Kubernetes/AIOps): Designing, deploying, and operating production Kubernetes control planes for AI GPU cloud workloads with an accent on topology-aware scheduling, automated remediation, and multi-tenant isolation. Focus on building GPU fault drain and rescheduling workflows, automating bare-metal tenant lifecycle management, and meeting cluster availability and job-completion SLOs.
Location: Remote within San Jose, California or Austin, Texas
Company
provides Bitcoin mining solutions and AI computational infrastructure, including data center, equipment management, operations, and cloud capabilities.
What you will do
- Design, deploy, and operate production Kubernetes clusters for GPU workloads ranging from 100 to 10,000 GPUs.
- Configure the Nvidia GPU Operator, device plugin, MIG, GPU time-slicing, and topology-aware scheduling across NVLink domains and network rails.
- Develop CRDs and integrate Kubernetes with Slurm, Ray, and Kubeflow for AI workload lifecycle management.
- Build multi-tenant isolation using namespaces, network policies, quotas, RBAC, and pod security standards.
- Automate BMaaS provisioning, tenant onboarding, lifecycle management, reclamation, and GPU node failure recovery.
- Operate monitoring and reliability systems with Prometheus, Grafana, Alertmanager, PagerDuty, SLIs/SLOs, runbook automation, and incident reviews.
Requirements
- 5+ years of Kubernetes operations experience, including at least 2 years managing GPU workloads on Kubernetes.
- Deep knowledge of the Nvidia GPU Operator, device plugin, GPU scheduling, topology-aware scheduling, and GPU resource management.
- Hands-on experience building multi-tenant Kubernetes platforms with strong isolation guarantees.
- Experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom systems.
- Proficiency in Terraform, Helm, GitOps workflows, ArgoCD or Flux, and programming in Go or Python for operator and CRD development.
- Strong SRE experience with SLI/SLO frameworks, incident management, capacity planning, monitoring at scale, and automated remediation.
Culture & Benefits
- Work on an AI-operated GPU cloud and infrastructure supporting AI computational workloads.
- Build an AIOps execution surface where remediation workflows can act safely without manual pager intervention.
- Apply a runbook-as-code approach so operational playbooks become executable platform workflows.
- Equal employment opportunities are provided in accordance with applicable country, state, and local laws.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →