5 дней назад
Storage Engineering Manager (AI Cloud)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Storage Engineering Manager (AI Cloud): Leading the storage engineering team that operates Ceph-based block, file, and object storage for large-scale AI workloads across multiple datacenters with an accent on team growth, reliability, capacity, and delivery. Focus on building mature on-call and incident processes, driving fleet-wide Ceph upgrades and object-storage launches, and creating repeatable storage operations for new datacenters.
Location: Hybrid from the Helsinki or London office, or fully remote within the EU.
Company
is building a full-stack AI cloud spanning datacenters, hardware, and a cloud platform for AI teams.
What you will do
- Lead the Storage team, including 1:1s, coaching, career development, performance management, compensation input, hiring, and onboarding.
- Set team structure, ownership areas, priorities, quarterly plans, milestones, and delivery expectations across storage services.
- Establish reliable operating practices including on-call rotations, escalation policies, runbooks, change management, postmortems, and written handovers.
- Own storage durability, availability, capacity, performance, upgrade, hardware, security, and compliance outcomes.
- Represent Storage to Infrastructure, Cloud, Platform, AI/ML, customer-facing, procurement, and data-center operations stakeholders.
- Drive projects including the customer launch of Object Storage, fleet-wide Ceph upgrades, shared filesystem strategy, VM volume features, and repeatable storage deployment for new datacenters.
Requirements
- 3+ years managing infrastructure or software engineers, including hiring, performance management, and leading teams through growth or change.
- 5+ years of hands-on production experience with distributed storage, ideally Ceph, including upgrades, recovery, rebalancing, capacity expansion, hardware failures, and incident response.
- Strong understanding of block, file, and object storage trade-offs, including RBD, CephFS, NFS, S3, replication, erasure coding, latency, throughput, and storage networking.
- Deep Linux systems and performance experience with NVMe, filesystems, MTU, bonding, and performance analysis.
- Experience designing or improving humane on-call and incident-management processes, with strong written communication skills.
- Ability to earn engineering trust through sufficient technical depth and turn ambiguity into actionable plans.
Nice to have
- Ceph RGW multisite, S3-compatible object storage, and customer-facing storage products.
- High-performance or parallel filesystems such as Lustre, VAST, DDN, WEKA, GPUDirect Storage, RDMA, or NVMe-oF.
- Storage for GPU clusters, ML training or inference, virtualization, QEMU/KVM, Kubernetes, CSI, or Rook.
- Multi-datacenter storage operations, repeatable site deployment, vendor management, or hardware and software evaluations.
- Storage security and compliance, including encryption at rest, data sanitization, ISO 27001, or SOC 2.
Culture & Benefits
- Work alongside engineers, researchers, and partners across the AI ecosystem.
- Collaborate in an international environment with 40+ nationalities.
- Receive cash and equity compensation along with local benefits.
- Join a growing infrastructure organization opening new datacenter sites regularly.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →