Staff ML Performance Engineer (Inference Optimisation)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Staff ML Performance Engineer (Inference Optimisation): Optimising ML inference for edge accelerators and GPUs to enable production-ready autonomous driving systems with an accent on large transformer-based models on low-power devices. Focus on profiling the full inference stack, implementing custom kernels, and ensuring reliability on in-vehicle compute.
Location: Hybrid (London, UK)
Company
is a leading developer of Embodied AI technology creating mapless and hardware-agnostic AI products for autonomous driving.
What you will do
- Profile and eliminate bottlenecks across the full inference stack, including model graphs and kernel execution.
- Implement optimizations in compilers, runtimes, and kernels, such as operator fusion and quantization.
- Develop benchmarking and regression testing frameworks to validate performance across different devices.
- Optimize targets for NVIDIA Orin/Thor and Qualcomm platforms.
- Collaborate with model developers to align architecture and training with on-device performance needs.
- Define technical roadmaps and elevate performance engineering standards across the organization.
Requirements
- Proven experience improving performance in production systems with strict latency, memory, and power constraints.
- Proficiency with toolchains such as TensorRT, CUDA, Qualcomm QNN, Triton, or OpenCL.
- Ability to work across levels of abstraction, from high-level model behavior to low-level kernel execution.
- Strong software engineering fundamentals in debugging, profiling, and writing maintainable code.
- Must be based in or able to work from the London office (Hybrid).
Nice to have
- Experience with embedded/edge deployment and benchmarking on real hardware.
- Expertise with NVIDIA and Qualcomm SoCs and their performance tools.
- Proficiency in Python and C++.
- Experience mentoring engineers or driving technical direction in fast-paced teams.
Culture & Benefits
- Hybrid working policy combining office collaboration with home-based work.
- Inclusive and diverse work environment that values new perspectives.
- Opportunity to work on groundbreaking Embodied AI for autonomous vehicles.
- Fast-paced environment focused on solving complex, high-impact problems.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →