2 дня назад
Helix AI Engineer, Training Performance (CUDA)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Helix AI Engineer, Training Performance (CUDA): Improving distributed training for 100B+ parameter models across 100k+ GPUs with an accent on GPU kernel optimization, accelerator evaluation, and large-scale training reliability. Focus on designing high-performance kernels and model training strategies, eliminating I/O and communication bottlenecks, and building resilient systems for failures across massive clusters.
Location: San Jose, CA, United States
Salary: $200,000–$400,000 annually
Company
is an AI robotics company developing autonomous general-purpose humanoid robots for home and commercial applications.
What you will do
- Optimize distributed training performance for 100B+ parameter models across 100k+ GPUs.
- Influence accelerator selection, cluster topology, scheduling, hardware procurement, and model co-design decisions.
- Write and optimize custom Triton and CUDA kernels, and contribute to kernel compilers such as Triton and Gluon.
- Build monitoring, regression detection, benchmarking, and root-cause analysis tooling for large-scale training jobs.
- Optimize data pipelines, checkpointing, fault tolerance, elastic restart, and model/data parallelism strategies.
- Evaluate AMD, TPU, SRAM-based ASIC, and other emerging accelerators through proof-of-concept ports and benchmarks.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer or Electrical Engineering, or a related field.
- 3+ years of AI performance engineering experience, including leadership of large-scale performance improvement projects.
- Deep understanding of GPU architecture, memory bandwidth, compute-bound and memory-bound operations, and occupancy.
- Experience with Nsight Systems/Compute, PyTorch Profiler, HTA, or similar profiling tools.
- Knowledge of NCCL, RDMA, NVLink, InfiniBand/RoCE, and topology-aware placement.
- Strong Python and CUDA/C++ skills, with experience debugging large-scale performance regressions, instability, and hardware-efficiency metrics such as MFU/HFU.
Nice to have
- Experience with heterogeneous or multi-datacenter training and cross-cluster orchestration.
- Open-source contributions to PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, or related ML systems projects.
- Exposure to AMD GPUs, TPU, Trainium, Inferentia, custom silicon, or heterogeneous fleet management.
Culture & Benefits
- Work on autonomous humanoid robots designed for global deployment.
- Collaborate across AI research, systems engineering, hardware, and infrastructure disciplines.
- Full-time employment with annual base salary of $200,000–$400,000.
- Total compensation may include additional components and benefits depending on the role.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 дня назад
Performance Engineer (Inference, Training & GPU)
200 000 - 300 000$
4 дня назад
Inference Engineer (AI)
195 000 - 285 000$
4 дня назад
Research Scientist (Embodied AI)
250 000 - 350 000$
4 дня назад
Systems Engineer (AI/HPC)
160 000 - 320 000$
4 дня назад
Systems/GPU Engineer (AI)
160 000 - 320 000$
5 дней назад
Software Engineer, Applied AI (AI)
150 000 - 170 000$