Назад
Company hidden
25 дней назад

Senior GPU Cloud Storage Solutions Expert (SRE SME) (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/Malaysia
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior GPU Cloud Storage Solutions Expert (SRE SME) (AI) (WEKA/VAST Data/Ceph/Lustre): Deploying and operating parallel storage systems and high-speed storage fabrics for AI and GPU cluster workloads with an accent on performance, multi-tenant isolation, and GPU Direct Storage. Focus on designing storage architectures for checkpoint I/O and inference, diagnosing latency and throughput issues, and converting storage incidents into predictive signals and executable runbooks.

Location: Singapore, SG / Penang, MY

Company

hirify.global is a technology company providing Bitcoin mining solutions, AI cloud capabilities, ASIC chips, mining rigs, and HPC datacenter operations.

What you will do

  • Deploy and operate parallel and distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre.
  • Design storage architectures for AI workloads such as checkpoint I/O bursts, sequential dataset reads, and inference KV cache.
  • Implement multi-tenant storage isolation with quotas, quality of service, access controls, and GPU Direct Storage.
  • Manage NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX orchestration.
  • Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest.
  • Build storage observability, fault-prediction signals, runbook-as-code automation, capacity plans, migration procedures, and disaster recovery processes.

Requirements

  • 5+ years of enterprise or HPC storage operations experience, including at least 2 years supporting AI/ML workloads.
  • Hands-on experience deploying and operating at least two of WEKA, VAST Data, Ceph, and DDN/Lustre.
  • Strong understanding of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers.
  • Experience with high-performance storage networking, GPU Direct Storage, RDMA-based data transfer, and multi-tenant storage QoS.
  • Strong Linux systems knowledge, including kernel tuning, filesystem internals, and block device management.
  • Experience with storage performance benchmarking and with turning storage telemetry and operational procedures into automated remediation.

Culture & Benefits

  • Inclusive environment valuing authenticity and diverse perspectives.
  • Startup spirit within a fast-growing technology company.
  • Autonomy, personal accountability, and opportunities to contribute to new systems and processes.
  • Training, mentoring, learning opportunities, and welfare benefits.
  • Opportunity to contribute to digital asset and AI infrastructure projects.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →