Назад
Company hidden
15 часов назад

Senior Staff Engineer (AI Data Path & Storage)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Staff Engineer (AI Data Path & Storage): Building high-performance storage and data movement architectures for real-time AI inference with an accent on NVIDIA NIXL, GPU–storage I/O paths, distributed storage, and KV cache management. Focus on optimizing latency-sensitive inference systems, integrating GPUDirect Storage and RDMA, and solving complex bottlenecks across GPU, networking, and storage subsystems.

Location: Remote in California with a hybrid work arrangement

Company

hirify.global develops the Infinia Data Intelligence Platform and advanced storage infrastructure for AI inference and large-scale distributed workloads.

What you will do

  • Lead the design and implementation of high-performance data movement pipelines using NVIDIA NIXL across GPU, CPU, and storage tiers.
  • Integrate hirify.global Infinia with GPU-accelerated inference platforms for large-scale, real-time AI workloads.
  • Optimize GPU-to-storage I/O paths using GPUDirect Storage, RDMA, InfiniBand, and NVMe-over-Fabrics.
  • Design multi-tier storage architectures and advanced KV cache strategies for latency, throughput, scalability, and data locality.
  • Partner with AI/ML teams to optimize PyTorch and TensorFlow inference performance and establish production benchmarking frameworks.
  • Mentor engineers and provide technical leadership across architecture, observability, automation, reliability, and performance initiatives.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 12+ years of experience in storage systems, distributed systems, or performance engineering.
  • Deep expertise in distributed storage architectures, Linux I/O, filesystem internals, storage protocols, NVMe, SSD optimization, and high-performance environments.
  • Hands-on experience with RDMA, InfiniBand, GPU computing, CPU–GPU data movement, and latency-sensitive production systems.
  • Proficiency in Python and/or C/C++, including advanced debugging, profiling, and performance tuning.
  • Must be based in California and work in a hybrid arrangement.

Nice to have

  • Experience with NVIDIA NIXL or comparable data movement frameworks and GPUDirect Storage.
  • Experience with AI inference systems, LLM serving, KV cache optimization, RAG pipelines, and vector search ecosystems.
  • Background in HPC or hyperscale distributed environments.
  • Experience with caching, memory tiering, data locality, and disaggregated compute and storage architectures.

Culture & Benefits

  • Hands-on work on GPU-native storage layers for AI inference.
  • Opportunity to build next-generation distributed AI infrastructure using NIXL and Infinia.
  • Focus on performance breakthroughs in real-time LLM inference at scale.
  • Technical leadership and cross-functional collaboration across storage, networking, GPU, and AI/ML engineering.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →