Назад
Company hidden
13 дней назад

Machine Learning Engineer, Performance Tooling (AI)

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer, Performance Tooling (AI) (ML Compilation): Building and extending an end-to-end ML compiler for deploying autonomous-driving models on NVIDIA TensorRT and Qualcomm QNN targets with an accent on quantization, graph lowering, and accuracy and latency validation. Focus on designing compiler passes, partitioning graphs across SoCs, preserving precision through decomposition, and scaling compilation infrastructure across architectures.

Location: London, United Kingdom or Sunnyvale, California, USA

Company

hirify.global builds autonomous driving technology that runs on real vehicles.

What you will do

  • Own the ML compilation pipeline from model checkpoint to deployable bundles for NVIDIA TensorRT and Qualcomm QNN targets.
  • Design and implement compiler passes covering capture, decomposition, precision assignment, legalisation, and graph partitioning.
  • Build accuracy and latency gates, regression suites, and benchmarking infrastructure for compiler changes.
  • Scale compilation infrastructure across model architectures, hardware platforms, and system-on-chip targets.
  • Partner with model and training teams on model compilability and set technical direction for compiler engineering.
  • Mentor others on compiler design and help shape the compilation stack at Staff level.

Requirements

  • Experience building or owning significant parts of an ML compilation or graph-lowering pipeline.
  • Deep experience with quantisation in compilation, including precision typing, post-training quantisation integration, and debugging accuracy loss caused by compiler transformations.
  • Strong Python skills and experience building and testing compiler infrastructure in production codebases.
  • Proficiency with at least one of MLIR, ONNX, TensorRT, Qualcomm QNN, or PyTorch graph capture/export.
  • Experience with multi-target compilation or graph partitioning across hardware backends.
  • Ability to assess correctness, accuracy, and performance trade-offs across compiler stages; C++ experience is a plus.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →