Назад
Company hidden
3 дня назад

Senior Platform Reliability Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Australia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Platform Reliability Engineer (AI): Operating GPU compute fleets, exabyte-scale storage, shared platform services, and observability infrastructure with an accent on production reliability, performance analysis, and guarded automation. Focus on diagnosing internals-level faults, building self-healing remediation, leading major incident recovery, and maintaining service levels across large-scale AI infrastructure.

Location: Based in Australia or Singapore, with travel to Australian AI Factory sites as required.

Company

hirify.global Technologies develops and operates energy-efficient AI infrastructure, including the hirify.global AI Cloud and AI FactoryOS platforms, across Asia Pacific.

What you will do

  • Operate the multi-tenant control and management plane, including tenancy, quota, and access controls.
  • Run exabyte-scale distributed filesystems, high-performance storage, and S3-compatible object storage to defined service levels.
  • Operate the production GPU fleet, including node health, fault handling, firmware and driver baselines, remediation, and hardware replacement workflows.
  • Build guarded automation, remediation tooling, provisioning frameworks, and CI/CD processes that support self-healing operations.
  • Diagnose and tune performance across Linux, networking, software-defined storage, RDMA, GPU Direct Storage, RoCE, and InfiniBand data paths.
  • Lead major-incident recovery, production-readiness reviews, patching and vulnerability remediation, vendor escalations, runbook development, and senior-engineer mentoring.

Requirements

  • 8+ years of infrastructure, systems, or platform engineering experience, including ownership of production storage and shared infrastructure services in a 24/7 environment.
  • Extensive experience with scale-out or parallel storage such as VAST, WEKA, Ceph, Lustre, GPFS, or NetApp.
  • Expert-level Linux knowledge covering storage and filesystem internals, kernel and driver behavior, networking, memory, I/O, and performance analysis.
  • Experience operating virtualization platforms, observability infrastructure, Kubernetes, infrastructure-as-code, GitOps, and operational automation using Python, Go, or Bash.
  • Experience with major-incident response, on-call escalation, post-incident reviews, vendor escalation, runbooks, least privilege, secrets management, certificates, and audited privileged operations.
  • Must be based in Australia or Singapore and able to travel to Australian AI Factory sites as required.

Nice to have

  • GPU or HPC infrastructure operations experience, including distributed training and inference environments.
  • Experience with multi-tenant cloud, service-provider, colocation, GPU-enabled Kubernetes, or bare-metal provisioning environments.
  • Knowledge of hardware lifecycle management, DPUs, SmartNICs, InfiniBand, RoCE, policy-as-code, software supply-chain controls, ISO 27001, or SOC 2.
  • Bachelor's degree in computer science, engineering, or a related discipline, or equivalent experience and training.

Culture & Benefits

  • Permanent full-time employment.
  • Participation in a shared after-hours escalation roster within a 24/7 operations function.
  • High degree of autonomy, broad technical direction, and direct access to decision makers.
  • Opportunity to work on large-scale, sustainable AI infrastructure and GPU systems.
  • Inclusive workplace welcoming candidates from diverse backgrounds.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →