Назад
Company hidden
2 дня назад

Staff Storage Platform Engineer (AI Storage)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Europe
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Storage Platform Engineer (AI Storage) (Distributed AI Storage): Designing and operating the AI storage layer for large-scale GPU infrastructure across edge and core deployments, with an accent on distributed inference, training, checkpointing, and model artifact delivery. Focus on optimizing throughput, latency, data locality, KV-cache persistence, and high-performance GPU-to-storage data paths using technologies such as Weka, VAST Data, NVMe-oF, RDMA, RoCE, and GPU Direct Storage.

Location: Europe

Company

Radian Arc provides an infrastructure-as-a-service platform for running cloud gaming, artificial intelligence, and machine learning applications inside telecom carrier networks. The platform provides GPU-based edge computing for low-latency services and 5G investment monetization.

What you will do

  • Design, build, and operate scalable AI storage architectures across edge and core deployments.
  • Optimize storage for distributed inference, fine-tuning, training, dataset ingestion, checkpointing, model artifacts, and KV-cache persistence.
  • Architect and integrate hyperconverged, local NVMe, and disaggregated storage platforms, including StorPool, VAST Data, and Weka.
  • Integrate block, object, and shared file storage with Kubernetes, orchestration systems, and CSI drivers.
  • Engineer high-performance data paths using RDMA, RoCE, GPU Direct Storage, SPDK, and NVMe-oF.
  • Lead performance investigations, incident response, reliability improvements, capacity planning, and end-to-end production delivery.

Requirements

  • Strong hands-on experience designing and operating distributed storage systems for high-performance computing and large-scale AI workloads.
  • Deep knowledge of the Linux storage and I/O stack, storage hardware, NVMe devices, storage fabrics, filesystems, and object and block storage.
  • Practical experience with Weka Data Platform, large-scale storage clusters, Kubernetes storage integrations such as CSI, and GPU-accelerated workloads.
  • Experience designing storage for distributed inference or training, including dataset distribution, checkpointing, model artifacts, and KV-cache storage.
  • Strong automation skills with Python and/or Bash, plus experience building reusable operational tooling.
  • Ability to lead complex cross-functional initiatives, set architectural direction, troubleshoot cross-layer issues, and mentor engineers without formal management authority.

Nice to have

  • Experience owning architecture and direct implementation in lean or fast-scaling environments.
  • Experience with StorPool, VAST Data, MinioFS, Rook Ceph, distributed file systems, and S3-compatible object storage.

Culture & Benefits

  • Attractive compensation package based on expertise and experience.
  • Friendly, internationally diverse, flexible, and hybrid-friendly work environment.
  • Opportunity to join a fast-growing scale-up focused on AI, cloud gaming, and edge infrastructure.
  • Career growth opportunities in a rapidly evolving infrastructure organization.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →