9 дней назад
Senior SRE & Automation Engineer (Customer-facing)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior SRE & Automation Engineer (Customer-facing) (GPU cloud/Kubernetes): Operating a customer-facing GPU cloud service and production Kubernetes clusters for workloads spanning 100–10,000 GPUs with an accent on tenant isolation, topology-aware scheduling, bare-metal provisioning, and service reliability. Focus on automating GPU fault remediation, building self-service observability, defining customer SLAs/SLOs, and turning incident runbooks into safe autonomous workflows.
Location: Singapore, SG / Penang, MY
Company
provides Bitcoin mining solutions, ASIC mining hardware, datacenter infrastructure, and AI cloud capabilities for customers worldwide.
What you will do
- Own end-to-end reliability of the customer-facing GPU cloud service, including availability, job completion, provisioning latency, and tenant experience.
- Operate production Kubernetes clusters for GPU workloads at scale, including Nvidia GPU Operator, device plugins, MIG, time-slicing, and multi-tenant allocation.
- Implement topology-aware scheduling using GPU locality, NVLink domains, and network rail affinity.
- Automate customer and tenant lifecycle management, bare-metal provisioning, isolation, quota management, handoff, and reclamation.
- Define customer-facing SLIs, SLOs, and SLAs; manage incidents, customer communications, post-incident reviews, and capacity planning.
- Build observability, infrastructure-as-code, and GitOps workflows while preparing the control plane for automated remediation and self-healing operations.
Requirements
- 5+ years of experience in SRE or cloud operations, including at least 2 years operating GPU workloads at scale.
- Deep Kubernetes operations experience and strong knowledge of GPU workload management, scheduling, and resource isolation.
- Experience building multi-tenant cloud platforms with strong isolation guarantees and operating customer-facing services against SLAs and SLOs.
- Hands-on experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom tooling.
- Proficiency with Terraform, Helm, GitOps workflows, Prometheus, Grafana, and large-scale alerting.
- Strong programming skills in Go or Python, plus experience with incident management, error budgets, capacity planning, and runbook automation.
Nice to have
- Experience developing AIOps control-plane workflows and automated remediation.
- Experience turning SRE playbooks into executable runbooks and platform workflows.
Culture & Benefits
- Inclusive environment that values authenticity and diverse perspectives.
- Startup-style environment within a fast-growing digital asset and AI cloud company.
- Autonomy, personal accountability, rapid growth, and learning opportunities.
- Training, mentoring, welfare benefits, and opportunities to contribute to new systems and projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →