Назад
Company hidden
6 дней назад

Machine Learning Performance Engineer, Training

200 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Performance Engineer, Training (Machine Learning/GPUs): Building and optimizing large-scale machine learning training systems across data ingestion, distributed execution, GPU kernels, and infrastructure with an accent on throughput, scalability, and hardware utilization. Focus on benchmarking heterogeneous compute platforms, optimizing multi-GPU and multi-node training, and solving memory, communication, numerical performance, and fault-tolerance challenges.

Location: New York, NY; hybrid working opportunities

Annual base salary: $200,000, plus eligibility for a discretionary bonus

Company

Quantitative trading firm developing high-performance electronic trading infrastructure, machine learning systems, and supporting business services.

What you will do

  • Benchmark machine learning training workloads across CPUs, GPUs, and other accelerator platforms.
  • Design and optimize distributed training strategies, including data, tensor, pipeline, and model parallelism.
  • Improve end-to-end training efficiency across data loading, memory management, model execution, checkpointing, and recovery.
  • Develop GPU kernels and performance-critical framework components using specialized libraries, compilers, and execution techniques.
  • Apply mixed precision, gradient accumulation, activation checkpointing, operator fusion, and memory-efficient optimization methods.
  • Collaborate with ML researchers, quantitative researchers, HPC engineers, systems engineers, and hardware specialists to build efficient training infrastructure.

Requirements

  • 3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments.
  • Deep knowledge of PyTorch or JAX, including execution models, compilation, autograd, and distributed training.
  • Strong Python and C++ programming skills, with experience developing performance-critical systems.
  • Experience developing and optimizing GPU kernels with CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries.
  • Strong understanding of GPU architecture, high-performance networking, storage, and accelerator interconnects such as InfiniBand, RDMA, NVLink, and NVSwitch.
  • Experience with distributed-training technologies and performance-analysis tools such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, Nsight Systems, Nsight Compute, or PyTorch Profiler.

Nice to have

  • Experience optimizing transformer, time-series, reinforcement-learning, or other computationally intensive models.
  • Experience with Kubernetes, Slurm, Ray, fault-tolerant distributed training, large-scale checkpointing, reproducibility, or GPU-cluster observability.
  • Experience with specialized accelerators, custom hardware, or machine learning compiler technologies.

Culture & Benefits

  • Hybrid work opportunities in a collaborative, results-oriented environment.
  • Generous paid time off and regional savings and financial wellness plans.
  • Free breakfast, lunch, and snacks, plus in-office wellness experiences and wellness expense reimbursement.
  • Volunteer opportunities, charitable giving, social events, and celebrations.
  • Workshops and continuous learning opportunities in a welcoming workplace without unnecessary hierarchy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →