2 дня назад
ML Performance Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Performance Engineer (AI): Optimising inference for transformer-based autonomous-driving models across edge accelerators, GPUs, compilers, runtimes and kernels with an accent on latency, memory, bandwidth, power and cost constraints. Focus on profiling model graphs, implementing compiler and kernel optimisations, building performance regression benchmarks, and deploying models to NVIDIA and Qualcomm edge platforms.
Location: London, United Kingdom; hybrid working model with in-person collaboration and remote work.
Company
builds an AI platform for autonomous driving, using embodied intelligence and scalable deployment to enable vehicles to learn from real-world experience.
What you will do
- Profile inference performance across model graphs, compiler and runtime behaviour, kernel execution, and memory movement.
- Implement and validate optimisations in compilers, runtimes, and kernels, including fusion, scheduling, quantisation-aware performance, and custom kernels.
- Build benchmarking and regression tests across models, devices, and software releases.
- Optimise ML inference for edge targets including NVIDIA Orin/Thor and Qualcomm platforms.
- Collaborate with model, platform, and deployment engineers on performance trade-offs and deployment-aware decisions.
- Contribute to tooling, documentation, and technical discussions around ML performance.
Requirements
- Experience improving production or production-adjacent systems under latency, memory, bandwidth, power, thermal, or cost constraints.
- Hands-on experience with at least one relevant stack, such as TensorRT, CUDA, Qualcomm QNN, Triton, or OpenCL.
- Ability to work from high-level model behaviour through to kernel- and runtime-level execution.
- Strong software engineering fundamentals in debugging, profiling, testing, and maintainable code.
- Clear communication and effective collaboration across ML, systems, and deployment teams.
- Availability to work in a hybrid setup in London, United Kingdom.
Nice to have
- Experience deploying or benchmarking ML models on embedded or edge devices.
- Familiarity with NVIDIA or Qualcomm SoCs and performance tooling.
- Proficiency in Python and C++.
- Experience supporting other engineers or contributing to technical direction within a small team.
Culture & Benefits
- Hybrid work with core hours, office collaboration, and hands-on work in vehicle workshops and labs.
- Market-benchmarked salaries, meaningful equity, and comprehensive location-dependent benefits.
- Relocation support and visa sponsorship where applicable.
- Learning and development budgets for training, conferences, and professional growth.
- Health insurance, dental coverage, enhanced parental leave, retirement or pension benefits where applicable, therapy access, wellbeing partnerships, and team socials.
- An evolving environment with significant ownership in shaping how the organisation works.
Hiring process
- Initial call or recruiter screen.
- Deep-dive technical interviews covering programming, system design, and domain-specific topics, taking approximately three hours in total.
- Final interview focused on mission and values alignment.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →