17 часов назад
AI and HPC Systems Performance Engineer
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI and HPC Systems Performance Engineer (AI/GPU/HPC): Building, deploying, benchmarking, and optimizing AI training and inference workloads across GPU-accelerated HPE infrastructure with an accent on distributed systems, telemetry, and performance analysis. Focus on tuning multi-GPU and multi-node environments, troubleshooting complex AI stacks, and developing automation, observability tools, and reference architectures.
Location: Hybrid, with an average requirement of two days per week from an HPE office
Company
delivers edge-to-cloud infrastructure and high-performance computing solutions for complex, data-intensive workloads.
What you will do
- Install, configure, and optimize GPU servers, storage, high-speed networking, and AI software stacks on HPE platforms.
- Benchmark and characterize AI and machine learning training, inference, LLM, multimodal, RAG, and distributed workloads.
- Analyze telemetry, metrics, logs, traces, and profiling data to identify bottlenecks and improve performance and scalability.
- Develop automation scripts, deployment frameworks, infrastructure-as-code solutions, and observability tools for AI platforms.
- Collaborate with customers, partners, GPU vendors, and internal engineering teams to troubleshoot and optimize AI solutions.
- Evaluate emerging AI technologies, author technical reports and reference architectures, and provide technical leadership and mentoring.
Requirements
- Typically 8+ years of experience in engineering or related technical fields.
- Strong Linux system administration and command-line experience across enterprise distributions.
- Experience with AI/ML frameworks and workloads, including PyTorch, JAX, model training, inference, benchmarking, and optimization.
- Experience with GPU-accelerated systems, CUDA, NCCL, distributed GPU environments, and multi-node AI clusters.
- Experience with high-performance networking, including InfiniBand, RDMA, RoCE, and Mellanox/NVIDIA networking solutions.
- Proficiency in Python, Bash, Go, C++, or a similar programming or scripting language; experience with Docker, Kubernetes, profiling, tracing, and observability tools.
Nice to have
- MS, ME, MTech, or PhD in computer science, computer engineering, electrical engineering, data science, artificial intelligence, or a related discipline.
- Experience with HPE platforms, AI Factory architectures, enterprise AI infrastructure, and large-scale foundation model workloads.
- Experience with Weka, Lustre, BeeGFS, GPFS, MLPerf, or other high-performance storage and benchmarking technologies.
- Experience developing reference architectures, technical papers, benchmark studies, or performance guidance.
Culture & Benefits
- Hybrid work structure with flexibility to manage work and personal needs.
- Health and wellbeing benefits supporting employees and their families.
- Personal and professional development programs aligned with career goals.
- Inclusive environment that values varied backgrounds and individual uniqueness.
- Opportunities to work with globally distributed teams and technology partners.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
2 дня назад
HPC / AI Software Infrastructure Lead (E)
151 100 - 256 900$
4 дня назад
Staff AI Infrastructure Engineer (AI)
241 000 - 331 000$
4 дня назад
AI Infrastructure Engineer (AI)
3 дня назад
Senior Software Engineer (AI)
165 450 - 259 750$
5 дней назад
AI Systems Software Engineer - Neuromorphic Computing
128 880 - 211 200$
3 дня назад