15 часов назад
Senior Staff Engineer (AI Data Path & Storage)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Staff Engineer (AI Data Path & Storage): Building high-performance storage and data movement architectures for real-time AI inference with an accent on NVIDIA NIXL, GPU–storage I/O paths, distributed storage, and KV cache management. Focus on optimizing latency-sensitive inference systems, integrating GPUDirect Storage and RDMA, and solving complex bottlenecks across GPU, networking, and storage subsystems.
Location: Remote in California with a hybrid work arrangement
Company
develops the Infinia Data Intelligence Platform and advanced storage infrastructure for AI inference and large-scale distributed workloads.
What you will do
- Lead the design and implementation of high-performance data movement pipelines using NVIDIA NIXL across GPU, CPU, and storage tiers.
- Integrate Infinia with GPU-accelerated inference platforms for large-scale, real-time AI workloads.
- Optimize GPU-to-storage I/O paths using GPUDirect Storage, RDMA, InfiniBand, and NVMe-over-Fabrics.
- Design multi-tier storage architectures and advanced KV cache strategies for latency, throughput, scalability, and data locality.
- Partner with AI/ML teams to optimize PyTorch and TensorFlow inference performance and establish production benchmarking frameworks.
- Mentor engineers and provide technical leadership across architecture, observability, automation, reliability, and performance initiatives.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- 12+ years of experience in storage systems, distributed systems, or performance engineering.
- Deep expertise in distributed storage architectures, Linux I/O, filesystem internals, storage protocols, NVMe, SSD optimization, and high-performance environments.
- Hands-on experience with RDMA, InfiniBand, GPU computing, CPU–GPU data movement, and latency-sensitive production systems.
- Proficiency in Python and/or C/C++, including advanced debugging, profiling, and performance tuning.
- Must be based in California and work in a hybrid arrangement.
Nice to have
- Experience with NVIDIA NIXL or comparable data movement frameworks and GPUDirect Storage.
- Experience with AI inference systems, LLM serving, KV cache optimization, RAG pipelines, and vector search ecosystems.
- Background in HPC or hyperscale distributed environments.
- Experience with caching, memory tiering, data locality, and disaggregated compute and storage architectures.
Culture & Benefits
- Hands-on work on GPU-native storage layers for AI inference.
- Opportunity to build next-generation distributed AI infrastructure using NIXL and Infinia.
- Focus on performance breakthroughs in real-time LLM inference at scale.
- Technical leadership and cross-functional collaboration across storage, networking, GPU, and AI/ML engineering.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 часов назад
Senior Solutions Architect (AI Infrastructure)
12 часов назад
Senior Principal AI System Architect
194 425 - 322 092$
16 часов назад
Senior NPU Architect (AI)
180 000 - 240 000$
14 часов назад
Principal Network Architect (AI Hardware)
250 000 - 300 000$
Databricks
3 дня назад
Senior Forward Deployed Engineer - Senior Architect (Data & AI)
18 часов назад