Назад
Company hidden
2 дня назад

AI Storage Solutions Expert

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Storage Solutions Expert (AI Storage): Deploying and operating high-performance distributed storage for AI training and inference across US data centers with an accent on checkpoint I/O, multi-tenant isolation, GPU Direct Storage, and storage telemetry. Focus on designing storage architectures for GB200-class GPU clusters, diagnosing performance faults, and turning incidents into predictive AIOps signals and executable remediation.

Location: Remote within San Jose, California or Austin, Texas

Company

hirify.global Technologies Group develops Bitcoin mining infrastructure and AI computational infrastructure, including GPU cloud, data centers, and cloud capabilities for AI workloads.

What you will do

  • Deploy and operate parallel and distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre.
  • Design storage architectures for AI checkpoint I/O bursts, sequential dataset reads, and inference KV cache workloads.
  • Implement multi-tenant isolation with QoS, quotas, access controls, and GPU Direct Storage data paths.
  • Deploy and manage NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX orchestration.
  • Benchmark and tune storage performance using fio, IOR, and mdtest, while maintaining failure-mode runbooks.
  • Plan capacity for GPU cluster growth and manage firmware, migrations, disaster recovery, and storage operations.

Requirements

  • 5+ years of enterprise or HPC storage operations experience, including at least 2 years supporting AI/ML workloads.
  • Hands-on experience deploying and operating at least two of WEKA, VAST Data, Ceph, and DDN/Lustre.
  • Strong understanding of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers.
  • Experience with high-performance storage networking, NFS over RDMA, NVMe-oF, GPU Direct Storage, and RDMA-based data transfer.
  • Proficiency in storage benchmarking and tuning, plus strong Linux knowledge covering kernel tuning, filesystem internals, and block devices.
  • Experience with multi-tenant storage isolation and QoS, storage telemetry, anomaly detection, and executable runbooks.

Culture & Benefits

  • Work on an AI-operated GPU cloud and storage infrastructure for large-scale training and inference.
  • Partner with the platform team to define storage-fault predictors, telemetry signals, incident labels, and false-positive tolerances.
  • Convert storage incidents and standard operating procedures into automated, agent-executable remediation.
  • Contribute to the storage design of Nvidia GB200-class clusters and improve storage incident MTTR.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →