Назад
Company hidden
12 дней назад

Sr. GPU Cloud Storage Solutions Expert (SRE SME)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Sr. GPU Cloud Storage Solutions Expert (SRE SME) (AI/GPU cloud storage): Deploying and operating high-performance parallel storage for AI training and inference across US data centers with an accent on distributed storage, GPU Direct Storage, RDMA networking, and multi-tenant isolation. Focus on instrumenting storage telemetry, building storage-fault prediction signals, tuning performance for GB200-class clusters, and converting incident procedures into executable automation.

Location: Remote within San Jose, CA or Austin, TX

Company

hirify.global builds AI computational infrastructure and Bitcoin mining solutions, including GPU cloud capabilities and data centers across multiple countries.

What you will do

  • Deploy and operate parallel and distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre.
  • Design storage architectures for AI checkpoint I/O bursts, sequential dataset reads, and inference KV cache workloads.
  • Implement multi-tenant isolation, quotas, QoS, access controls, and GPU Direct Storage data paths.
  • Manage high-performance storage networking including NFS over RDMA, NVMe-oF, storage fabrics, and Nvidia CMX.
  • Benchmark and tune storage performance using fio, IOR, and mdtest; manage capacity planning, firmware, migration, and disaster recovery.
  • Turn storage incidents into telemetry-driven prediction and runbook-as-code automation for the AIOps platform.

Requirements

  • 5+ years of enterprise or HPC storage operations experience, including at least 2 years supporting AI/ML workloads.
  • Hands-on experience deploying and operating at least two of WEKA, VAST Data, Ceph, or DDN/Lustre.
  • Strong understanding of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers.
  • Experience with NFS over RDMA, NVMe-oF, GPU Direct Storage, and RDMA-based data transfer.
  • Proficiency in storage benchmarking and tuning with fio, IOR, and mdtest, plus strong Linux systems knowledge.
  • Experience with multi-tenant storage isolation and QoS, storage telemetry, anomaly detection, and executable runbooks.

Culture & Benefits

  • Work on an AI-operated GPU cloud and high-performance storage infrastructure.
  • Partner with the platform team to define storage-fault prediction signals, labels, and false-positive tolerances.
  • Build observability and a baseline predictor for the top storage-fault classes.
  • Help deliver Nvidia GB200-class clusters based on the storage architecture.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →