Назад
4 дня назад

Member of Technical Staff Storage Infrastructure

150 000 - 300 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
senior
Страна
US
vacancy_detail.hirify_telegram_tooltipВакансия из Telegram канала -

Мэтч & Сопровод

Покажет вашу совместимость и напишет письмо

Описание вакансии

TL;DR
Member of Technical Staff Storage Infrastructure (Distributed Storage): Designing and operating storage systems for datasets, checkpoints, artifacts, and research workflows with an accent on parallel filesystems, object storage, NVMe caching, and reliability. Focus on benchmarking throughput and latency, automating provisioning and lifecycle management, and testing replication, backup, and recovery procedures.

Member of Technical Staff Storage Infrastructure

Company

Prime Intellect

Conditions

4 days agoSalary: 150K - 300K

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will design and operate storage systems for datasets, checkpoints, artifacts, and research workflows. You will tune storage platforms, benchmark performance, automate operations, test recovery procedures, and improve reliability, security, and observability.

Requirements

  • 3+ years building or operating production distributed storage systems
  • Experience with Lustre, BeeGFS, Ceph, GPFS, or another parallel, distributed filesystem or object storage platform
  • Linux administration and performance troubleshooting skills
  • Experience automating infrastructure operations with Python, Go, Bash, or similar languages
  • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
  • Knowledge of block, file, object storage, NVMe, filesystem tuning, I/O profiling, authentication, authorization, and encryption

Responsibilities

  • Design and operate storage architectures for datasets, checkpoints, inference artifacts, and research workflows
  • Deploy and tune parallel filesystems, object storage, and NVMe caching
  • Benchmark throughput, latency, metadata performance, and concurrent access
  • Build storage provisioning, capacity planning, lifecycle management, and operational automation
  • Design and test replication, recovery, backup, and failure-handling procedures
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks

Benefits

  • Equity incentives

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Текст вакансии взят без изменений

Источник -