Назад
Company hidden
17 часов назад

AI and HPC Systems Performance Engineer

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
c1
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI and HPC Systems Performance Engineer (AI/GPU/HPC): Building, deploying, benchmarking, and optimizing AI training and inference workloads across GPU-accelerated HPE infrastructure with an accent on distributed systems, telemetry, and performance analysis. Focus on tuning multi-GPU and multi-node environments, troubleshooting complex AI stacks, and developing automation, observability tools, and reference architectures.

Location: Hybrid, with an average requirement of two days per week from an HPE office

Company

hirify.global delivers edge-to-cloud infrastructure and high-performance computing solutions for complex, data-intensive workloads.

What you will do

  • Install, configure, and optimize GPU servers, storage, high-speed networking, and AI software stacks on HPE platforms.
  • Benchmark and characterize AI and machine learning training, inference, LLM, multimodal, RAG, and distributed workloads.
  • Analyze telemetry, metrics, logs, traces, and profiling data to identify bottlenecks and improve performance and scalability.
  • Develop automation scripts, deployment frameworks, infrastructure-as-code solutions, and observability tools for AI platforms.
  • Collaborate with customers, partners, GPU vendors, and internal engineering teams to troubleshoot and optimize AI solutions.
  • Evaluate emerging AI technologies, author technical reports and reference architectures, and provide technical leadership and mentoring.

Requirements

  • Typically 8+ years of experience in engineering or related technical fields.
  • Strong Linux system administration and command-line experience across enterprise distributions.
  • Experience with AI/ML frameworks and workloads, including PyTorch, JAX, model training, inference, benchmarking, and optimization.
  • Experience with GPU-accelerated systems, CUDA, NCCL, distributed GPU environments, and multi-node AI clusters.
  • Experience with high-performance networking, including InfiniBand, RDMA, RoCE, and Mellanox/NVIDIA networking solutions.
  • Proficiency in Python, Bash, Go, C++, or a similar programming or scripting language; experience with Docker, Kubernetes, profiling, tracing, and observability tools.

Nice to have

  • MS, ME, MTech, or PhD in computer science, computer engineering, electrical engineering, data science, artificial intelligence, or a related discipline.
  • Experience with HPE platforms, AI Factory architectures, enterprise AI infrastructure, and large-scale foundation model workloads.
  • Experience with Weka, Lustre, BeeGFS, GPFS, MLPerf, or other high-performance storage and benchmarking technologies.
  • Experience developing reference architectures, technical papers, benchmark studies, or performance guidance.

Culture & Benefits

  • Hybrid work structure with flexibility to manage work and personal needs.
  • Health and wellbeing benefits supporting employees and their families.
  • Personal and professional development programs aligned with career goals.
  • Inclusive environment that values varied backgrounds and individual uniqueness.
  • Opportunities to work with globally distributed teams and technology partners.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →