Назад
Company hidden
2 дня назад

Senior AI Storage Infrastructure Engineer (AI)

Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Singapore
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Senior AI Storage Infrastructure Engineer (Distributed Storage): Designing and maintaining high-performance storage systems for AI cloud infrastructure with an accent on CSI drivers, GPUDirect Storage (GDS), and RDMA optimizations. Focus on eliminating CPU bottlenecks, optimizing IOPS/latency for massive datasets, and scaling production HPC storage environments.

Location: Singapore, SG

Company

hirify.global is a world-leading technology company specializing in Bitcoin mining and AI cloud services.

What you will do

  • Design, deploy, and maintain CSI drivers for high-performance parallel file systems such as Weka, Lustre, DAOS, and VAST.
  • Implement GPUDirect Storage (GDS) to enable direct memory access (DMA) between NVMe drives and GPU memory.
  • Develop local NVMe caching strategies for rapid loading of massive model weights and datasets during distributed training.
  • Optimize IOPS, throughput, and latency profiles across the entire containerized storage stack.
  • Collaborate with GPU Systems & Fabric teams to ensure storage layer optimization for RDMA and high-speed interconnects (InfiniBand, RoCE).
  • Define storage policies, quota management, and multi-tenancy isolation strategies within Kubernetes.

Requirements

  • 5+ years of experience in distributed storage systems and high-performance file systems.
  • Deep expertise in the Kubernetes CSI paradigm, including building volume plugins and storage operators.
  • Strong hands-on experience with Linux OS block/file I/O and kernel-level performance tuning.
  • Familiarity with high-throughput networking protocols (RDMA, InfiniBand, RoCE).
  • Proven track record of operating and scaling large-scale storage environments in production or HPC settings.
  • Experience with infrastructure automation tools like Terraform, Ansible, and CI/CD pipelines.

Culture & Benefits

  • Inclusive and respectful environment with an open workspace and startup spirit.
  • Opportunity to network with industrial pioneers and contribute to the future of the digital asset industry.
  • High degree of personal accountability, autonomy, and fast growth opportunities.
  • Attractive welfare benefits and developmental opportunities, including training and mentoring.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →