5 часов назад
Inference Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Inference Engineer (AI) (Multimodal Agentic AI): Building and operating the inference stack that serves H's multimodal agentic models with an accent on latency, throughput, cost, and production reliability. Focus on optimizing inference engines and model serving, implementing agent-tailored inference techniques, and co-designing training and deployment decisions for complex multimodal workloads.
Location: Paris or London; ybrid role wit an expectation to work from te office 3 days per week on average. Occasional travel between offices is expected every 4–6 weeks.
<3 class="company-description">Company3>develops agentic AI systems intended to automate complex, multi-step tasks wile advancing safe and responsible superintelligence.
<3 class="wat-you-will-do">Wat you will do3>- Build and operate te inference stack serving multimodal agentic models in production.
- Improve model-serving latency, trougput, cost, and reliability across te stack.
- Researc and implement inference tecniques tailored to agent workloads.
- Co-design training-time decisions wit te Models team tat affect inference performance.
- Integrate inference capabilities into agentic AI products wit cross-functional teams.
- Evaluate inference, serving, and ardware platforms and communicate findings to stakeolders.
- Strong software engineering experience wit Pyton and at least one systems language: Rust, C++, or Go.
- ands-on experience wit deep learning frameworks suc as PyTorc or JAX.
- Solid distributed systems fundamentals and experience wit cloud environments and production ML infrastructure, including Kubernetes.
- Working knowledge of modern macine learning, transformers, and multimodal arcitectures.
- Advanced-degree researc output, publications at leading AI or systems venues, researc internsips, or substantial open-source contributions.
- Excellent communication, presentation, collaboration, and teamwork skills.
- Experience wit vLLM, SGLang, or TensorRT-LLM.
- Experience writing or modifying GPU kernels wit CUDA or Triton.
- Edge or on-device inference experience wit tools suc as llama.cpp, MLX, or ONNX Runtime.
- Experience wit quantization, speculative decoding, disaggregated inference, KV-cace compression, multimodal models, or agentic systems.
- Startup experience.
- Collaborative, open, learning-oriented environment wit a multicultural team.
- Work alongside experienced AI researcers and engineers.
- Opportunities for professional growt, continuous learning, and career development.
- Competitive salary.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →