Назад
Company hidden
2 дня назад

Helix AI Engineer, Training Performance (CUDA)

200 000 - 400 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Helix AI Engineer, Training Performance (CUDA): Improving distributed training for 100B+ parameter models across 100k+ GPUs with an accent on GPU kernel optimization, accelerator evaluation, and large-scale training reliability. Focus on designing high-performance kernels and model training strategies, eliminating I/O and communication bottlenecks, and building resilient systems for failures across massive clusters.

Location: San Jose, CA, United States

Salary: $200,000–$400,000 annually

Company

hirify.global is an AI robotics company developing autonomous general-purpose humanoid robots for home and commercial applications.

What you will do

  • Optimize distributed training performance for 100B+ parameter models across 100k+ GPUs.
  • Influence accelerator selection, cluster topology, scheduling, hardware procurement, and model co-design decisions.
  • Write and optimize custom Triton and CUDA kernels, and contribute to kernel compilers such as Triton and Gluon.
  • Build monitoring, regression detection, benchmarking, and root-cause analysis tooling for large-scale training jobs.
  • Optimize data pipelines, checkpointing, fault tolerance, elastic restart, and model/data parallelism strategies.
  • Evaluate AMD, TPU, SRAM-based ASIC, and other emerging accelerators through proof-of-concept ports and benchmarks.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Computer or Electrical Engineering, or a related field.
  • 3+ years of AI performance engineering experience, including leadership of large-scale performance improvement projects.
  • Deep understanding of GPU architecture, memory bandwidth, compute-bound and memory-bound operations, and occupancy.
  • Experience with Nsight Systems/Compute, PyTorch Profiler, HTA, or similar profiling tools.
  • Knowledge of NCCL, RDMA, NVLink, InfiniBand/RoCE, and topology-aware placement.
  • Strong Python and CUDA/C++ skills, with experience debugging large-scale performance regressions, instability, and hardware-efficiency metrics such as MFU/HFU.

Nice to have

  • Experience with heterogeneous or multi-datacenter training and cross-cluster orchestration.
  • Open-source contributions to PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, or related ML systems projects.
  • Exposure to AMD GPUs, TPU, Trainium, Inferentia, custom silicon, or heterogeneous fleet management.

Culture & Benefits

  • Work on autonomous humanoid robots designed for global deployment.
  • Collaborate across AI research, systems engineering, hardware, and infrastructure disciplines.
  • Full-time employment with annual base salary of $200,000–$400,000.
  • Total compensation may include additional components and benefits depending on the role.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →