7 дней назад
Staff Engineer – L4 (AI Storage)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff Engineer – L4 (AI Storage): Resolving critical customer escalations and distributed-system issues in Infinia AI storage environments with an accent on root-cause analysis, performance tuning, and system resiliency. Focus on leading incident response, using AI-powered diagnostics and automation to reduce MTTR, and translating technical findings into product improvements and executive-ready RCAs.
Location: Pune office, India; on-site
Company
develops AI and multi-cloud data management and high-performance storage solutions for demanding data-intensive environments.
What you will do
- Own critical enterprise and hyperscale customer escalations, performing root-cause analysis and defining mitigation strategies.
- Lead Infinia incident response, war rooms, live incident bridges, and cross-functional collaboration with Engineering, QA, Field, Product Management, and customer-facing teams.
- Investigate complex issues across storage, distributed systems, protocols, applications, and performance, and propose product improvements or workarounds.
- Use AI-powered debugging, log analysis, system pattern recognition, observability, and automation to accelerate resolution and reduce MTTR.
- Develop runbooks, performance-tuning guides, RCA documentation, executive-ready summaries, and business-impact statements.
- Deliver training to customer support and field engineering and contribute to post-mortems and strategic-account briefings.
Requirements
- On-site work at the Pune office is required.
- 8+ years of experience in enterprise storage, distributed systems, or cloud infrastructure support or engineering.
- Deep knowledge of file systems and storage technologies, including S3, POSIX, NFS, storage performance, and Linux kernel internals.
- Strong scripting and coding skills in Python, Go, and C++, with hands-on Linux troubleshooting experience.
- Proven debugging ability at system, protocol, and application levels using tools such as strace, tcpdump, and perf.
- Exceptional communication and executive reporting skills; participation in an on-call rotation is required.
Nice to have
- Experience with , VAST, Weka, or similar scale-out file systems.
- Exposure to RDMA, NVMe-oF, or high-performance networking stacks.
- Familiarity with Prometheus, Grafana, ELK, or OpenTelemetry.
- Knowledge of replication, consistency models, and data integrity mechanisms.
- Experience with Sovereign AI, LLM model training environments, autonomous-system data architectures, or AI-based diagnostic tooling.
Culture & Benefits
- Hands-on work within the Infinia Core engineering team on critical AI storage infrastructure.
- Participation in an on-call rotation for after-hours support as needed.
- Structured Infinia training, labs, architecture deep dives, and customer-escalation shadowing.
- Opportunity to improve internal tooling, automation, documentation, reliability, and support diagnostics.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →