Назад
Company hidden
3 дня назад

Inference Engineer (AI)

140 000 - 200 000GBP
Тип работы
fulltime
Английский
b2
Страна
UK
Релокация
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Engineer (AI): Building and operating realtime voice-model inference systems at scale with an accent on model acceleration, low-latency serving, and high-performance GPU systems. Focus on optimizing quantization, batching, caching, distributed inference, and production reliability for thousands of concurrent connections.

Location: United Kingdom. Candidates must already have the legal right to work in the UK. Visa sponsorship is not available; future relocation to the San Francisco Bay Area may include U.S. visa and relocation support subject to business needs and work authorization requirements.

Base salary: £140,000–£200,000 per year, plus equity and benefits.

Company

hirify.global is an AI research lab developing realtime voice models, inference infrastructure, APIs, and products for consumer-facing applications.

What you will do

  • Optimize realtime model serving and sub-second multimodal inference at scale.
  • Apply quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
  • Build high-performance inference systems using C++, CUDA, Rust, or highly optimized Python.
  • Scale multi-GPU and multi-node inference with Kubernetes, Ray, custom load balancing, and reliable production infrastructure.
  • Take models from research through containerization, serving optimization, deployment, and production reliability.
  • Profile systems and improve latency, throughput, and stability for thousands of concurrent connections.

Requirements

  • Deep understanding of modern serving frameworks and inference optimization techniques, such as vLLM or TRT-LLM.
  • Proficiency in C++, CUDA, Rust, or highly optimized Python, with experience profiling NVIDIA GPU workloads.
  • Experience with distributed systems, Kubernetes, Ray, load balancing, and multi-GPU or multi-node inference.
  • Evidence of building non-trivial systems, contributing to open source, or producing deep technical work.
  • Ability to own the full lifecycle from research model to reliable production service.
  • PhD in computer science, physics, or mathematics, or equivalent practical experience building backend or ML systems.

Nice to have

  • Open-source contributions to major inference engines.
  • Deep-dive technical write-ups or other public technical work.

Culture & Benefits

  • Work on state-of-the-art realtime voice models used by large consumer-facing AI applications.
  • High autonomy in clarifying ambiguous problems through benchmarks and prototypes.
  • Performance, latency, reliability, and shipped impact are treated as first-class priorities.
  • Flat structure, fast iterations, and minimal process overhead.
  • Compensation includes equity and benefits in addition to base salary.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →