35 минут назад
Engineer, Storage and Data Protection (AI/HPC)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Engineer, Storage and Data Protection (AI/HPC): Providing enterprise-level operational support and engineering distributed storage and data protection systems for production AI/HPC workloads with an accent on storage performance, automation, and resilience. Focus on troubleshooting storage, Linux, networking, and I/O bottlenecks, optimizing data movement workflows, and evaluating scalable storage architectures.
Location: Hybrid in Gurugram, Haryana, Bangalore, or Hyderabad, India
Company
's Managed Services – MS Infrastructure department provides enterprise infrastructure operations and support for managed services customers.
What you will do
- Provide operational support for incident, problem, and change management activities.
- Administer and tune distributed and parallel filesystems, including Lustre, GPFS, BeeGFS, Ceph, Weka, and Vast.
- Optimize storage performance, throughput, metadata operations, and data locality for AI training and inference workloads.
- Build and maintain automation for provisioning, monitoring, alerting, quota management, and storage lifecycle operations.
- Support data ingest, replication, caching, tiering, archiving, and data protection workflows.
- Investigate complex storage, Linux, network, and I/O issues; coordinate with technical teams, customers, vendors, infrastructure, platform, and research teams.
- Evaluate storage architectures for scalability, resilience, and cost efficiency, maintain documentation, and participate in an on-call rotation.
Requirements
- 5+ years of experience with HPC, AI infrastructure, or large-scale storage engineering.
- Bachelor's degree or equivalent experience in Information Systems or a related field.
- Strong Linux systems administration experience and hands-on experience configuring, managing, and tuning distributed or parallel filesystems.
- Knowledge of HPC schedulers such as Slurm and/or container platforms such as Kubernetes, plus familiarity with InfiniBand or RDMA.
- Understanding of replication, backup strategies, and disaster recovery in HPC environments, including machine learning or data science workflows.
- Managed Services or consulting experience, strong customer service, problem-solving, and communication skills.
Nice to have
- Experience supporting GPU clusters, AI/ML workflows, multi-petabyte environments, caching architectures, or storage isolation in multi-tenant systems.
- Familiarity with object storage such as S3, MinIO, or Ceph Object Gateway.
- Experience with Terraform, Ansible, Helm, GitOps, Prometheus, or Grafana.
- Scripting or programming experience with Python and Bash.
- Related storage certifications.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
Senior Enterprise Platform Support Engineer – AI & Cloud
16 часов назад
Senior Cloud Infrastructure Engineer (Cybersecurity)
2 дня назад
Platform Specialist - PLS (Storage)
100 000$
2 дня назад
Operations Software Analyst II (Control-M)
14 часов назад
Senior IT Operations Engineer
80 000 - 115 000$
12 часов назад