7 дней назад
Machine Learning Engineer (Synthetic Data)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Machine Learning Engineer (Synthetic Data) (Generative World Models): Building and scaling GAIA-class world models and high-throughput synthetic-data generation infrastructure for autonomous driving with an accent on controllable video generation, 3D geometry, and downstream model training. Focus on optimising multi-GPU inference, diagnosing calibration and controllability failures, and measuring the impact of synthetic data on driving models and safety-critical scenarios.
Location: Based in the London office, United Kingdom; hybrid working policy with time in the office and from home.
Company
develops embodied AI software and foundation models for mapless, hardware-agnostic autonomous driving systems.
What you will do
- Post-train and iterate GAIA-class world models for rig transfer, pose transfer, and conditioning based on geometry, calibration, and actions.
- Own the synthetic-data generation loop from configuration and large-scale GPU inference to training-ready artefacts with clear model and settings lineage.
- Integrate synthetic data into behaviour-cloning, reward-model, and reinforcement-learning training, including binarisation, mix ratios, quality filters, and impact experiments.
- Diagnose geometry, calibration, controllability, novel-view-synthesis, odometry, and camera-layout failures that affect generated video quality.
- Improve inference throughput, valid-generation rates, and self-service workflows using techniques such as distillation, KV caching, shortcut methods, and reduced step counts.
- Expand synthetic-data coverage to new vehicle platforms and safety-critical scenarios such as Emergency Lane Keeping and Automatic Emergency Braking.
Requirements
- 4+ years of applied machine learning or research engineering experience, including training and shipping neural networks.
- Strong Python and PyTorch or equivalent skills, with experience debugging GPU training and reading model code.
- Hands-on experience with video, generative, or world models, including diffusion, flow matching, autoregressive video, novel-view synthesis, or neural rendering.
- Working knowledge of cameras and 3D geometry, including multi-camera rigs, intrinsics, extrinsics, warps, and reprojection.
- Experience taking generated or simulated data into downstream models and measuring impact through mixes, ablations, and failure analysis.
- Experience operating reliable multi-GPU generation or training workflows and large video artefacts at scale.
Nice to have
- Experience with controllable generation, video diffusion or flow models, distillation, few-step sampling, or KV caching.
- Background in autonomous vehicles, robotics, simulation, or multi-sensor driving data.
- Experience with Flyte, Ray, Spark-style jobs, dataset lineage, or training-mix configuration.
- Experience with reward models, offline reinforcement learning, closed-loop driving-policy evaluation, cloud GPU fleets, or distributed training.
Culture & Benefits
- Hybrid work combining office and home working, with core working hours and flexibility to determine the schedule with the team.
- High-trust, high-autonomy environment focused on generative simulation and real-world autonomous-driving impact.
- Inclusive and collaborative culture with close work across machine-learning research, infrastructure, and driving-model teams.
- Opportunity to work with GAIA-scale video generation, camera transfer, pose control, and vehicle-training infrastructure.
Hiring process
- provides an inclusive interview experience and can accommodate candidates who require adjustments.
- Application materials may be used to pre-fill the application form.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →