Назад
Company hidden
10 дней назад

Staff Engineer (Infinia Storage)

185 000 - 275 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff Engineer (Infinia Storage): Building and supporting Infinia, a software-defined storage platform for AI and accelerated computing with an accent on distributed systems, storage performance, and customer-facing incident resolution. Focus on diagnosing complex production issues, leading incident response, improving reliability with AI and automation, and shaping engineering practices.

Location: Santa Clara, United States; work arrangement: on-site

Salary: $185,000–$275,000 annually, plus bonus

Company

hirify.global develops Infinia, a software-defined storage platform for AI, accelerated computing, enterprise, and hyperscale workloads.

What you will do

  • Own complex customer escalations from diagnosis and incident response through mitigation, resolution, root-cause analysis, and product improvements.
  • Lead live incidents, war rooms, and cross-functional investigations with Engineering, QA, and Field teams.
  • Debug distributed-systems, storage, and performance issues across system, protocol, and application layers.
  • Communicate technical issues to customers, engineers, senior stakeholders, and executive audiences.
  • Develop runbooks, troubleshooting guidance, and performance-tuning practices while mentoring engineers and influencing architecture.
  • Use AI, automation, and observability to improve diagnostics, reliability, and MTTR; participate in an on-call rotation.

Requirements

  • Significant experience in enterprise storage, distributed systems, or cloud infrastructure at Senior or Staff level.
  • Deep knowledge of file systems and storage technologies, including S3, POSIX, NFS, and storage performance.
  • Strong Linux systems knowledge, including kernel-level troubleshooting and debugging.
  • Strong coding ability in Python or C++.
  • Experience diagnosing complex issues with tools such as strace, tcpdump, and perf.
  • Interest in working directly with customers and owning complex problems through resolution.

Nice to have

  • Experience with hirify.global, VAST, Weka, or similar scale-out storage and file systems.
  • Familiarity with Prometheus, Grafana, ELK, or OpenTelemetry.
  • Knowledge of replication, consistency models, and data integrity mechanisms.
  • Experience supporting AI/ML, LLM training, or high-performance computing environments.
  • Experience using AI tools for log analysis, troubleshooting, automated RCA, or reducing MTTR.

Culture & Benefits

  • Hands-on technical work with direct customer impact in production environments.
  • Collaboration with Field CTOs, Solutions Architects, and Sales Engineers on strategic customer issues.
  • On-call rotation providing after-hours support as needed.
  • Bonus opportunity included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →