10 часов назад
Senior Lab Reliability Engineer (Storage)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Senior Lab Reliability Engineer (Storage): Maintaining production-quality reliability for VAST clusters and the infrastructure supporting a presales lab with an accent on storage operations, automation, and infrastructure as code. Focus on resolving complex issues across storage, networking, and compute, building diagnostic tooling, and validating customer-relevant scenarios.
Location: Remote - United States
Company
is building a presales platform that provides tooling, automation, and infrastructure for field engineering evaluations, demos, and internal enablement.
What you will do
- Own the operational reliability of VAST clusters in the lab, including health monitoring, upgrade planning, and issue resolution.
- Act as the technical escalation point for complex cluster issues and partner with VAST engineering on deeper investigations.
- Shape provisioning scripts, CLI utilities, monitoring dashboards, and internal lab tooling.
- Establish infrastructure-as-code and configuration-management standards using Ansible or equivalent tools.
- Reproduce and isolate difficult issues, producing diagnostic data and technical reports for engineering.
- Manage supporting infrastructure across virtualization, compute, networking, and storage while mentoring lab operations engineers.
Requirements
- 4+ years of experience in systems engineering, storage engineering, customer support engineering, or a related role.
- Hands-on operational experience with enterprise storage systems such as VAST, Pure, NetApp, Isilon, or Ceph.
- Strong Linux administration skills covering networking, storage, filesystems, systemd, and CLI tooling.
- Strong scripting or programming experience with Python, Bash, or similar tools.
- Operational experience with Docker, Kubernetes, infrastructure as code or configuration management, and virtualization platforms such as VMware vSphere, ESXi, or Proxmox.
- Solid networking, troubleshooting, technical documentation, and communication skills, with comfort working across time zones.
Nice to have
- Hands-on experience with clusters.
- Experience as a Customer Support Engineer or Reliability Engineer/SRE in storage or infrastructure.
- Familiarity with Grafana, Prometheus, Elasticsearch, or network switch administration.
- Experience with high-performance computing or ML/AI training workloads.
Culture & Benefits
- Remote engineering role supporting distributed team members.
- Collaboration with lab and platform engineers, pre-sales SEs, professional services, and engineering.
- Opportunity to mentor lab operations engineers and raise technical and operational standards.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →