Назад
обновлено 3 дня назад

Staff Storage Software Engineer (AI Infrastructure)

349 000 - 465 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Storage Software Engineer (AI Infrastructure): Building high-performance distributed storage infrastructure for AI clusters across object, block, and file protocols with an accent on petabyte-scale deployments, storage performance, and hardware-accelerated data paths. Focus on designing concurrent systems software, integrating NVMe/GPU-direct/DPU technologies, and solving reliability and performance challenges across complex production failure domains.

Location: Hybrid, with presence in the San Francisco or Bellevue office 4 days per week; Tuesday is the designated work-from-home day. San Jose office location is also listed.

Salary: $349,000–$465,000 annually for San Francisco/San Jose; $314,000–$419,000 annually for Bellevue.

Company

Lambda builds AI cloud infrastructure and high-performance compute, storage, and networking systems for researchers, enterprises, and hyperscalers.

What you will do

  • Set technical direction for storage software architecture across petabyte-scale deployments and lead architectural reviews.
  • Design, develop, and maintain high-performance storage systems across NFS, SMB, Lustre, NVMe-oF, iSCSI, and S3.
  • Build distributed systems for storage orchestration and integrate them with NVMe, GPU-direct, and DPU-accelerated hardware.
  • Own storage software delivery from requirements and design through deployment, monitoring, benchmarking, profiling, and capacity planning.
  • Partner with networking, control plane, Kubernetes, observability, compute, and fleet engineering teams to define and track storage SLOs and SLIs.
  • Optimize storage for AI workloads including checkpoint I/O, high-throughput dataset serving, and latency-sensitive inference pipelines.

Requirements

  • 10+ years of storage systems engineering experience, including 5+ years in a technical lead or Staff+ individual contributor role.
  • Production experience designing and operating multi-petabyte storage infrastructure in data center or cloud environments.
  • Strong proficiency in C, C++, Rust, or Go, with experience writing high-performance concurrent systems software.
  • Hands-on production expertise with at least two object, block, or file storage protocols.
  • Experience profiling and tuning throughput, latency, and IOPS using tools such as fio or elbencho.
  • Working knowledge of NVMe, NVMe-oF, RDMA, DPUs, physical data centers, storage reliability, incident response, and observability tools.

Nice to have

  • Experience with NVIDIA BlueField DPUs, SuperNICs, or GPUDirect Storage.
  • Production experience with Vast Data, Weka, NetApp, Lustre, or Ceph at scale.
  • Familiarity with CXL memory pooling, computational storage, ZNS SSDs, or open-source storage projects.

Culture & Benefits

  • Cash and equity compensation.
  • Health, dental, and vision coverage for employees and dependents.
  • Wellness and commuter stipends for select roles.
  • 401k plan with a 2% company match for USA employees.
  • Flexible paid time off.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →