5 дней назад
Sr SRE & Automation Engineer (Customer Facing)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Sr SRE & Automation Engineer (Customer Facing) (GPU cloud): Owns the reliability of a customer-facing GPU cloud service across tenant onboarding, Kubernetes workload execution, incident response, and recovery with an accent on GPU scheduling, multi-tenant isolation, and observability. Focus on automating GPU fault handling, bare-metal provisioning, customer-facing SLAs/SLOs, and AIOps remediation workflows for infrastructure scaling to 10,000 GPUs.
Location: Remote within San Jose, California or Austin, Texas
Company
provides Bitcoin mining infrastructure, AI computational infrastructure, and cloud capabilities for high-demand artificial intelligence workloads.
What you will do
- Own end-to-end reliability of a customer-facing GPU cloud service, including availability, job completion, provisioning latency, and tenant experience.
- Operate production Kubernetes clusters for GPU workloads at 100–10,000 GPUs, including NVIDIA GPU Operator, device plugins, MIG, time-slicing, and topology-aware scheduling.
- Build tenant lifecycle management with onboarding, quotas, isolation, RBAC, network policies, resource quotas, offboarding, and reclamation.
- Automate Bare-Metal-as-a-Service provisioning, tenant handoff, lifecycle management, and reclamation.
- Define customer-facing SLIs, SLOs, and SLAs; manage incidents, customer communications, post-incident reviews, capacity planning, and operational readiness.
- Develop observability and automation using Prometheus, Grafana, Alertmanager, PagerDuty, Terraform, Helm, and GitOps workflows.
Requirements
- 5+ years of experience in SRE or cloud operations, including at least 2 years operating GPU workloads at scale.
- Deep Kubernetes operations experience and expertise in GPU workload management, scheduling, and resource allocation.
- Experience building multi-tenant cloud platforms with strong isolation guarantees and operating customer-facing services against SLAs and SLOs.
- Hands-on experience with bare-metal server provisioning and lifecycle automation using Ironic, MAAS, or custom solutions.
- Proficiency in Terraform, Helm, GitOps, Prometheus, Grafana, and alerting at scale.
- Strong programming skills in Go or Python, plus experience with incident management, error budgets, capacity planning, AIOps, and executable runbooks.
Culture & Benefits
- Customer reliability is treated as a core product responsibility.
- Operational practices are designed to turn manual incident interventions into autonomous workflows.
- Customers receive self-service visibility into job health, quotas, and service status.
- Equal employment opportunities are provided in accordance with applicable laws.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →