4 часа назад
Storage-Focused SRE (AI Infrastructure)
170 000 - 205 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Storage-Focused SRE (AI Infrastructure) (Distributed Cloud Storage): Building and operating fault-tolerant block, file, and object storage systems for AI and HPC workloads with an accent on automation, self-healing infrastructure, and storage reliability. Focus on designing scalable storage backends, optimizing Linux I/O paths and file systems, and resolving complex incidents across high-performance NVMe- and SSD-backed infrastructure.
Location: Sunnyvale, California, United States; on-site
Salary: $170,000–$205,000 per year plus bonus and Restricted Stock Units
Company
builds vertically integrated, sustainable AI infrastructure spanning energy, data centers, cloud services, and compute.
What you will do
- Build automation and self-healing tools for distributed block, file, and object storage infrastructure.
- Drive reliability initiatives covering data replication, encryption, backup and restore, and failover mechanisms.
- Implement and maintain high-performance NVMe- and SSD-backed volumes for large-scale AI compute clusters.
- Operate user-facing storage services, improving availability, performance, and error-budget adherence.
- Investigate storage incidents using telemetry, logs, and performance profiling.
- Partner with storage, hardware, and kernel engineers to diagnose low-level I/O issues and design scalable storage backends.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- 5+ years of professional experience in Storage SRE, systems, or storage engineering.
- Hands-on experience operating enterprise storage platforms such as Pure Storage or EMC.
- Strong understanding of object, block, and file storage paradigms and managed storage services at scale.
- Proficiency in Go, Python, Java, or C, plus Infrastructure as Code and deployment tools such as Terraform, Ansible, or Puppet.
- Deep Linux internals knowledge, including I/O subsystems, memory management, and storage scheduling, with experience in storage protocols and container orchestration.
Nice to have
- Experience with Ceph, GlusterFS, OpenEBS, Vast, Lightbits, or other distributed storage systems.
- Contributions to open-source storage projects or the Linux storage stack.
- Experience with hybrid on-premises and cloud storage models.
Culture & Benefits
- Industry-competitive compensation with Restricted Stock Units.
- Health, vision, dental, HSA, life insurance, and disability coverage options.
- 401(k) with a 100% employer match up to 4% of salary.
- Paid parental leave, generous paid time off, and holidays.
- Tuition reimbursement, commuter benefits, cell phone reimbursement, and other wellbeing and legal benefits.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
7 часов назад
SRE Engineer II (Cloud/DevOps)
141 000 - 162 000$
6 часов назад
Senior Site Reliability Engineer (SRE)
170 000 - 196 000$
6 часов назад
Senior Site Reliability Engineer (AI Infrastructure)
215 000 - 275 000$
3 часа назад
Software Engineer, DevOps (Robotics)
115 000 - 170 000$
23 часа назад
Platform Engineer II (Cloud Infrastructure)
115 000 - 130 000$
2 часа назад
Senior Software Engineer, Site Reliability Engineering (AWS)
153 000 - 210 000$