3 дня назад
Machine Leaning Performance Engineer (Inference)
200 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Leaning Performance Engineer (Inference) (Machine Learning/Inference): Architecting and optimizing inference pipelines for high-performance quantitative trading systems with an accent on GPU kernels, heterogeneous hardware benchmarking, and microsecond-level latency. Focus on optimizing memory hierarchies, developing custom kernels, applying quantization and pruning, and deploying models across CPUs, GPUs, and FPGAs.
Location: New York, United States; hybrid working opportunities
Salary: $200,000–$300,000 annual base salary, plus an eligible discretionary bonus.
Company
is a quantitative trading firm developing high-performance electronic trading infrastructure, including low-latency systems, hardware acceleration, and machine learning platforms.
What you will do
- Evaluate inference platforms across CPUs, GPUs, and FPGAs to guide infrastructure deployment decisions.
- Analyze memory hierarchies, interconnects, and execution bottlenecks across the inference lifecycle.
- Design inference strategies that meet the thermal, power, and operational constraints of latency-sensitive trading hardware.
- Develop optimized GPU kernels and integrate specialized performance libraries.
- Optimize models through quantization, pruning, and distillation for compact memory usage, numerical stability, and real-time inference.
- Collaborate with ML researchers, HPC engineers, FPGA engineers, and datacenter engineers on production deployments.
Requirements
- 2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments.
- Deep expertise in lower-level ML framework development with PyTorch or JAX, plus strong Python and C++ skills.
- Strong understanding of mixed-precision computation and GPU microarchitecture, including SM execution, warp scheduling, and memory hierarchy optimization.
- Proven experience developing custom GPU kernels and using tools such as Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS, Nsight Systems, or Nsight Compute.
- Experience benchmarking inference performance across heterogeneous compute architectures using rigorous, data-driven methods.
Nice to have
- Practical experience optimizing inference workloads for specialized hardware ecosystems such as FPGAs and ASICs.
- Previous financial trading experience is not required.
Culture & Benefits
- Generous paid time off policies.
- Savings plans and financial wellness tools available in each region.
- Free breakfast, lunch, and snacks daily.
- In-office wellness experiences and reimbursement for selected wellness expenses.
- Company-sponsored sports teams, fitness events, volunteer opportunities, charitable giving, social events, and celebrations.
- Workshops and continuous learning opportunities in a collaborative, low-hierarchy environment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 дней назад
ML Engineer (Video Generation)
175 000 - 275 000$
3 дня назад
Member of Technical Staff – ML Systems & Inference (AI)
250 000 - 350 000$
5 дней назад
Staff Machine Learning Engineer (AI)
232 000 - 348 000$
4 дня назад
Machine Learning Engineer II (AI)
140 000 - 180 000$
4 дня назад
Machine Learning Engineer II (AI)
140 000 - 180 000$
5 дней назад
Machine Learning Engineer II (AI)
111 000 - 231 250$