Назад
Company hidden
обновлено 1 день назад

Staff ML Performance Engineer (Inference Optimisation)

Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/

TL;DR

Staff ML Performance Engineer (Inference Optimisation): Optimising ML inference for edge accelerators and GPUs to enable production-ready autonomous driving systems with an accent on large transformer-based models on low-power devices. Focus on profiling the full inference stack, implementing custom kernels, and ensuring reliability on in-vehicle compute.

Location: Hybrid (London, UK)

Company

hirify.global is a leading developer of Embodied AI technology creating mapless and hardware-agnostic AI products for autonomous driving.

What you will do

  • Profile and eliminate bottlenecks across the full inference stack, including model graphs and kernel execution.
  • Implement optimizations in compilers, runtimes, and kernels, such as operator fusion and quantization.
  • Develop benchmarking and regression testing frameworks to validate performance across different devices.
  • Optimize targets for NVIDIA Orin/Thor and Qualcomm platforms.
  • Collaborate with model developers to align architecture and training with on-device performance needs.
  • Define technical roadmaps and elevate performance engineering standards across the organization.

Requirements

  • Proven experience improving performance in production systems with strict latency, memory, and power constraints.
  • Proficiency with toolchains such as TensorRT, CUDA, Qualcomm QNN, Triton, or OpenCL.
  • Ability to work across levels of abstraction, from high-level model behavior to low-level kernel execution.
  • Strong software engineering fundamentals in debugging, profiling, and writing maintainable code.
  • Must be based in or able to work from the London office (Hybrid).

Nice to have

  • Experience with embedded/edge deployment and benchmarking on real hardware.
  • Expertise with NVIDIA and Qualcomm SoCs and their performance tools.
  • Proficiency in Python and C++.
  • Experience mentoring engineers or driving technical direction in fast-paced teams.

Culture & Benefits

  • Hybrid working policy combining office collaboration with home-based work.
  • Inclusive and diverse work environment that values new perspectives.
  • Opportunity to work on groundbreaking Embodied AI for autonomous vehicles.
  • Fast-paced environment focused on solving complex, high-impact problems.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →