Назад
Company hidden
5 дней назад

Sr SRE & Automation Engineer (Customer Facing)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr SRE & Automation Engineer (Customer Facing) (GPU cloud): Owns the reliability of a customer-facing GPU cloud service across tenant onboarding, Kubernetes workload execution, incident response, and recovery with an accent on GPU scheduling, multi-tenant isolation, and observability. Focus on automating GPU fault handling, bare-metal provisioning, customer-facing SLAs/SLOs, and AIOps remediation workflows for infrastructure scaling to 10,000 GPUs.

Location: Remote within San Jose, California or Austin, Texas

Company

hirify.global provides Bitcoin mining infrastructure, AI computational infrastructure, and cloud capabilities for high-demand artificial intelligence workloads.

What you will do

  • Own end-to-end reliability of a customer-facing GPU cloud service, including availability, job completion, provisioning latency, and tenant experience.
  • Operate production Kubernetes clusters for GPU workloads at 100–10,000 GPUs, including NVIDIA GPU Operator, device plugins, MIG, time-slicing, and topology-aware scheduling.
  • Build tenant lifecycle management with onboarding, quotas, isolation, RBAC, network policies, resource quotas, offboarding, and reclamation.
  • Automate Bare-Metal-as-a-Service provisioning, tenant handoff, lifecycle management, and reclamation.
  • Define customer-facing SLIs, SLOs, and SLAs; manage incidents, customer communications, post-incident reviews, capacity planning, and operational readiness.
  • Develop observability and automation using Prometheus, Grafana, Alertmanager, PagerDuty, Terraform, Helm, and GitOps workflows.

Requirements

  • 5+ years of experience in SRE or cloud operations, including at least 2 years operating GPU workloads at scale.
  • Deep Kubernetes operations experience and expertise in GPU workload management, scheduling, and resource allocation.
  • Experience building multi-tenant cloud platforms with strong isolation guarantees and operating customer-facing services against SLAs and SLOs.
  • Hands-on experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions.
  • Proficiency in Terraform, Helm, GitOps, Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python, plus experience with incident management, error budgets, capacity planning, AIOps, and executable runbooks.

Culture & Benefits

  • Customer reliability is treated as a core product responsibility.
  • Operational practices are designed to turn manual incident interventions into autonomous workflows.
  • Customers receive self-service visibility into job health, quotas, and service status.
  • Equal employment opportunities are provided in accordance with applicable laws.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →