Назад
Company hidden
4 часа назад

Inference Engineer (AI)

Формат работы
hybrid
Тип работы
fulltime
Английский
b2
Страна
France/UK
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Engineer (AI): Build and operate the inference stack serving multimodal agentic AI models with an accent on latency, throughput, cost optimization, and inference techniques tailored to agent workloads. Focus on designing and improving inference systems, collaborating with model and agent teams, and evaluating hardware platforms to power next-generation superintelligent AI.

Location

Hybrid in Paris or London, expected to be in the office 3 days a week on average with some travel between offices every 4-6 weeks.

Company

hirify.global pushes the boundaries of superintelligence with agentic AI, focusing on building safe and responsible AI agents that automate complex human tasks.

What you will do

  • Build and operate the inference stack serving multimodal agentic models.
  • Improve latency, throughput, and cost of model serving.
  • Research and implement inference techniques tailored to agent workloads.
  • Co-design with Models team on training-time decisions affecting inference.
  • Collaborate with cross-functional teams to integrate inference into AI products.
  • Evaluate inference, serving, and hardware platforms and communicate findings.
  • Stay current with advancements in inference, model serving, and accelerator technology.

Requirements

  • Location: Hybrid role in Paris or London with office presence required.
  • Strong software engineering skills with proficiency in Python and at least one systems language (Rust, C++, or Go).
  • Experience with deep learning frameworks (PyTorch, JAX) in industry settings.
  • Solid distributed systems fundamentals and experience with cloud and ML infrastructure (Kubernetes).
  • Working knowledge of modern ML including transformers and multimodal architectures.
  • Research engagement demonstrated by advanced degree, publications, internships, or open-source contributions.
  • Excellent communication and collaboration skills.

Nice to have

  • Startup experience.
  • Hands-on experience with inference frameworks (vLLM, SGLang, TensorRT-LLM).
  • Experience writing or modifying GPU kernels (CUDA, Triton).
  • Edge or on-device inference experience (llama.cpp, MLX, ONNX Runtime).
  • Experience with quantization, speculative decoding, disaggregated inference, or KV-cache compression.
  • Experience with multimodal models and/or agentic systems.

Culture & Benefits

  • Collaborate with a dynamic, multicultural team of world-class AI talent.
  • Competitive salary and opportunities for professional growth and continuous learning.
  • Work in a highly collaborative environment focused on advancing AI safely and responsibly.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →