Назад
Company hidden
3 дня назад

Machine Leaning Performance Engineer (Inference)

200 000 - 300 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Leaning Performance Engineer (Inference) (Machine Learning/Inference): Architecting and optimizing inference pipelines for high-performance quantitative trading systems with an accent on GPU kernels, heterogeneous hardware benchmarking, and microsecond-level latency. Focus on optimizing memory hierarchies, developing custom kernels, applying quantization and pruning, and deploying models across CPUs, GPUs, and FPGAs.

Location: New York, United States; hybrid working opportunities

Salary: $200,000–$300,000 annual base salary, plus an eligible discretionary bonus.

Company

hirify.global is a quantitative trading firm developing high-performance electronic trading infrastructure, including low-latency systems, hardware acceleration, and machine learning platforms.

What you will do

  • Evaluate inference platforms across CPUs, GPUs, and FPGAs to guide infrastructure deployment decisions.
  • Analyze memory hierarchies, interconnects, and execution bottlenecks across the inference lifecycle.
  • Design inference strategies that meet the thermal, power, and operational constraints of latency-sensitive trading hardware.
  • Develop optimized GPU kernels and integrate specialized performance libraries.
  • Optimize models through quantization, pruning, and distillation for compact memory usage, numerical stability, and real-time inference.
  • Collaborate with ML researchers, HPC engineers, FPGA engineers, and datacenter engineers on production deployments.

Requirements

  • 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments.
  • Deep expertise in lower-level ML framework development with PyTorch or JAX, plus strong Python and C++ skills.
  • Strong understanding of mixed-precision computation and GPU microarchitecture, including SM execution, warp scheduling, and memory hierarchy optimization.
  • Proven experience developing custom GPU kernels and using tools such as Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, or Nsight Compute.
  • Experience benchmarking inference performance across heterogeneous compute architectures using rigorous, data-driven methods.

Nice to have

  • Practical experience optimizing inference workloads for specialized hardware ecosystems such as FPGAs and ASICs.
  • Previous financial trading experience is not required.

Culture & Benefits

  • Generous paid time off policies.
  • Savings plans and financial wellness tools available in each region.
  • Free breakfast, lunch, and snacks daily.
  • In-office wellness experiences and reimbursement for selected wellness expenses.
  • Company-sponsored sports teams, fitness events, volunteer opportunities, charitable giving, social events, and celebrations.
  • Workshops and continuous learning opportunities in a collaborative, low-hierarchy environment.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →