Назад
Company hidden
9 дней назад

Senior SRE & Automation Engineer (Customer-facing)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior SRE & Automation Engineer (Customer-facing) (GPU cloud/Kubernetes): Operating a customer-facing GPU cloud service and production Kubernetes clusters for workloads spanning 100–10,000 GPUs with an accent on tenant isolation, topology-aware scheduling, bare-metal provisioning, and service reliability. Focus on automating GPU fault remediation, building self-service observability, defining customer SLAs/SLOs, and turning incident runbooks into safe autonomous workflows.

Location: Singapore, SG / Penang, MY

Company

hirify.global provides Bitcoin mining solutions, ASIC mining hardware, datacenter infrastructure, and AI cloud capabilities for customers worldwide.

What you will do

  • Own end-to-end reliability of the customer-facing GPU cloud service, including availability, job completion, provisioning latency, and tenant experience.
  • Operate production Kubernetes clusters for GPU workloads at scale, including Nvidia GPU Operator, device plugins, MIG, time-slicing, and multi-tenant allocation.
  • Implement topology-aware scheduling using GPU locality, NVLink domains, and network rail affinity.
  • Automate customer and tenant lifecycle management, bare-metal provisioning, isolation, quota management, handoff, and reclamation.
  • Define customer-facing SLIs, SLOs, and SLAs; manage incidents, customer communications, post-incident reviews, and capacity planning.
  • Build observability, infrastructure-as-code, and GitOps workflows while preparing the control plane for automated remediation and self-healing operations.

Requirements

  • 5+ years of experience in SRE or cloud operations, including at least 2 years operating GPU workloads at scale.
  • Deep Kubernetes operations experience and strong knowledge of GPU workload management, scheduling, and resource isolation.
  • Experience building multi-tenant cloud platforms with strong isolation guarantees and operating customer-facing services against SLAs and SLOs.
  • Hands-on experience with bare-metal provisioning and lifecycle automation using Ironic, MAAS, or custom tooling.
  • Proficiency with Terraform, Helm, GitOps workflows, Prometheus, Grafana, and large-scale alerting.
  • Strong programming skills in Go or Python, plus experience with incident management, error budgets, capacity planning, and runbook automation.

Nice to have

  • Experience developing AIOps control-plane workflows and automated remediation.
  • Experience turning SRE playbooks into executable runbooks and platform workflows.

Culture & Benefits

  • Inclusive environment that values authenticity and diverse perspectives.
  • Startup-style environment within a fast-growing digital asset and AI cloud company.
  • Autonomy, personal accountability, rapid growth, and learning opportunities.
  • Training, mentoring, welfare benefits, and opportunities to contribute to new systems and projects.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →