25 дней назад
Senior GPU Cloud Storage Solutions Expert (SRE SME) (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior GPU Cloud Storage Solutions Expert (SRE SME) (AI) (WEKA/VAST Data/Ceph/Lustre): Deploying and operating parallel storage systems and high-speed storage fabrics for AI and GPU cluster workloads with an accent on performance, multi-tenant isolation, and GPU Direct Storage. Focus on designing storage architectures for checkpoint I/O and inference, diagnosing latency and throughput issues, and converting storage incidents into predictive signals and executable runbooks.
Location: Singapore, SG / Penang, MY
Company
is a technology company providing Bitcoin mining solutions, AI cloud capabilities, ASIC chips, mining rigs, and HPC datacenter operations.
What you will do
- Deploy and operate parallel and distributed storage systems including WEKA, VAST Data, Ceph, and DDN/Lustre.
- Design storage architectures for AI workloads such as checkpoint I/O bursts, sequential dataset reads, and inference KV cache.
- Implement multi-tenant storage isolation with quotas, quality of service, access controls, and GPU Direct Storage.
- Manage NFS over RDMA, NVMe-oF, high-speed storage fabrics, and Nvidia CMX orchestration.
- Diagnose and tune storage performance using IOPS, throughput, latency profiling, fio, IOR, and mdtest.
- Build storage observability, fault-prediction signals, runbook-as-code automation, capacity plans, migration procedures, and disaster recovery processes.
Requirements
- 5+ years of enterprise or HPC storage operations experience, including at least 2 years supporting AI/ML workloads.
- Hands-on experience deploying and operating at least two of WEKA, VAST Data, Ceph, and DDN/Lustre.
- Strong understanding of AI training I/O patterns, including checkpoint frequency, dataset loading, and shuffle buffers.
- Experience with high-performance storage networking, GPU Direct Storage, RDMA-based data transfer, and multi-tenant storage QoS.
- Strong Linux systems knowledge, including kernel tuning, filesystem internals, and block device management.
- Experience with storage performance benchmarking and with turning storage telemetry and operational procedures into automated remediation.
Culture & Benefits
- Inclusive environment valuing authenticity and diverse perspectives.
- Startup spirit within a fast-growing technology company.
- Autonomy, personal accountability, and opportunities to contribute to new systems and processes.
- Training, mentoring, learning opportunities, and welfare benefits.
- Opportunity to contribute to digital asset and AI infrastructure projects.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
11 дней назад
Sr. Site Reliability Engineer (AI)
Nscale
4 дня назад
Senior Operational Engineer (AI)
14 дней назад
Senior SRE (Site Reliability Engineer) – Modernized Application Operations
145 000 - 170 000$
11 дней назад
Senior SRE (Kubernetes)
150 000 - 170 000$
13 дней назад
Senior Database Reliability Engineer (AI)
10 дней назад