10 дней назад
Storage Engineer (AI Infrastructure)
150 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Storage Engineer (AI Infrastructure) (Distributed Storage/GPUs): Building and operating high-throughput storage systems for frontier AI training, inference artifacts, datasets, and checkpoints with an accent on parallel filesystems, object storage, NVMe caching, durability, and availability. Focus on benchmarking and tuning storage under concurrent GPU workloads, automating capacity and lifecycle management, and designing recovery, replication, and failure-handling procedures.
Location: San Francisco or Remote
Salary: $150,000–$300,000 per year plus equity incentives
Company
is building an open superintelligence stack that provides frontier AI infrastructure for research teams, enterprises, and AI companies.
What you will do
- Design and operate storage architectures for training datasets, checkpoints, inference artifacts, and shared research workflows.
- Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads.
- Benchmark throughput, latency, metadata performance, and concurrent access using representative training and checkpoint workloads.
- Build provisioning, capacity planning, lifecycle management, and operational automation for storage services.
- Design and test replication, recovery, backup, and failure-handling procedures with defined durability and availability targets.
- Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices while collaborating with compute and networking teams.
Requirements
- 3+ years of experience building or operating production distributed storage systems.
- Hands-on experience with a parallel or distributed filesystem or object storage platform such as Lustre, BeeGFS, Ceph, or GPFS.
- Strong Linux administration and performance troubleshooting skills.
- Experience automating infrastructure operations with Python, Go, Bash, or similar languages.
- Understanding of storage failure modes, data integrity, consistency, replication, and recovery.
- Knowledge of storage semantics, NVMe/SSD performance, filesystem tuning, I/O profiling, storage networking, observability, access control, encryption, and secure data lifecycle management.
Nice to have
- Experience supporting large GPU training clusters and high-volume checkpoint workloads.
- Experience with S3-compatible object storage, data tiering, distributed caching, RDMA-enabled storage, or GPUDirect Storage.
- Experience with Kubernetes storage integrations or SLURM environments.
- Storage cost optimization experience or contributions to open-source storage systems.
Culture & Benefits
- Work directly with customers training foundation models and deploying large-scale inference infrastructure.
- Collaborate with an engineering team building open frontier AI infrastructure.
- Have direct impact on reliable, high-performance GPU infrastructure.
- Receive cash compensation of $150,000–$300,000 plus equity incentives.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 дней назад
Senior Cloud Infrastructure Engineer (AWS/GCP)
200 000 - 230 000$
12 дней назад
Forward Deployed Engineer (AI Connectivity)
192 010 - 274 300$
11 дней назад
Infrastructure Automation Engineer (Terraform)
100 000 - 180 000$
13 дней назад
Engineering Manager (AI)
200 000 - 300 000$
13 дней назад
Automation Infrastructure Engineer (AI)
170 000 - 200 000$
13 дней назад
AI Operations Manager (AI)
220 000 - 240 000$