Inference Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Location: United Kingdom. Candidates must already have the legal right to work in the UK. Visa sponsorship is not available; future relocation to the San Francisco Bay Area may include U.S. visa and relocation support subject to business needs and work authorization requirements.
Base salary: £140,000–£200,000 per year, plus equity and benefits.
Company
is an AI research lab developing realtime voice models, inference infrastructure, APIs, and products for consumer-facing applications.
What you will do
- Optimize realtime model serving and sub-second multimodal inference at scale.
- Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
- Build high-performance inference systems using C++, CUDA, Rust, or highly optimized Python.
- Scale multi-GPU and multi-node inference with Kubernetes, Ray, custom load balancing, and reliable production infrastructure.
- Take models from research through containerization, serving optimization, deployment, and production reliability.
- Profile systems and improve latency, throughput, and stability for thousands of concurrent connections.
Requirements
- Deep understanding of modern serving frameworks and inference optimization techniques, such as vLLM or TRT-LLM.
- Proficiency in C++, CUDA, Rust, or highly optimized Python, with experience profiling NVIDIA GPU workloads.
- Experience with distributed systems, Kubernetes, Ray, load balancing, and multi-GPU or multi-node inference.
- Evidence of building non-trivial systems, contributing to open source, or producing deep technical work.
- Ability to own the full lifecycle from research model to reliable production service.
- PhD in computer science, physics, or mathematics, or equivalent practical experience building backend or ML systems.
Nice to have
- Open-source contributions to major inference engines.
- Deep-dive technical write-ups or other public technical work.
Culture & Benefits
- Work on state-of-the-art realtime voice models used by large consumer-facing AI applications.
- High autonomy in clarifying ambiguous problems through benchmarks and prototypes.
- Performance, latency, reliability, and shipped impact are treated as first-class priorities.
- Flat structure, fast iterations, and minimal process overhead.
- Compensation includes equity and benefits in addition to base salary.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →