Назад
Company hidden
19 часов назад

Senior AI Storage Infrastructure Engineer (AI)

Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore/US/Norway +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior AI Storage Infrastructure Engineer (AI storage and Kubernetes): Building the high-performance data-delivery fabric for an AI-native NeoCloud with an accent on distributed storage, GPU-direct I/O, and low-latency access for large-scale training and inference workloads. Focus on designing CSI drivers, integrating GPUDirect Storage, optimizing NVMe caching and storage performance, and enabling RDMA-connected infrastructure.

Location: Remote within San Jose, California or Austin, Texas

Company

hirify.global develops Bitcoin mining solutions and AI computational and cloud infrastructure across multiple countries.

What you will do

  • Design, deploy, and maintain CSI drivers for high-performance parallel file systems such as Weka, Lustre, DAOS, and VAST.
  • Architect GPUDirect Storage integrations and direct NVMe-to-GPU memory data paths for AI workloads.
  • Develop local NVMe caching strategies for large model weights and datasets used in distributed training.
  • Optimize IOPS, throughput, and latency across the containerized storage stack.
  • Collaborate on RDMA, InfiniBand, and RoCE integration with GPU infrastructure.
  • Implement monitoring, alerting, storage policies, quotas, and Kubernetes multi-tenancy isolation; mentor engineers and lead architecture reviews.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 5+ years of experience with distributed storage systems and high-performance file systems, including POSIX compliance and file I/O semantics.
  • Deep expertise in Kubernetes CSI, volume plugins, and storage operators.
  • Strong Linux block and file I/O experience, including kernel-level performance tuning.
  • Experience with RDMA, InfiniBand, RoCE, production storage operations, and large-scale or HPC environments.
  • Experience with Terraform, Ansible, CI/CD pipelines, technical communication, and cross-functional architecture decisions.

Nice to have

  • Experience in high-velocity, high-growth engineering environments.

Culture & Benefits

  • Full-time employment.
  • Work remotely within the specified United States locations.
  • Work on AI infrastructure and large-scale distributed computing systems.
  • Collaborate across storage, GPU systems, fabric, and infrastructure engineering.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →