Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Compiler Engineer (AI/GPU): Building AGCO, an agentic compiler that explores, generates, compiles, verifies, and benchmarks LLM execution optimizations on real GPUs with an accent on compiler IRs, GPU systems, and LLM inference. Focus on designing optimization passes, code generation, verification loops, and search methods that improve execution performance across AMD and NVIDIA hardware.
Location: Paris, France; hybrid work with the ability to relocate to Paris and work closely with the team
Company
KOG builds a co-designed inference stack for real-time AI agents across model architecture, inference engines, compilers, and low-level GPU kernels.
What you will do
- Build AGCO, an agentic compiler for optimizing LLM execution across different GPUs and optimization targets.
- Design compiler and intermediate-representation systems for representing and transforming LLM computations.
- Develop optimization passes, lowering pipelines, and code generation.
- Create search methods that explore alternative implementations and execution strategies.
- Build verification and correctness checks for generated changes.
- Profile and optimize GPU execution, memory, parallelism, communication, and LLM inference performance.
Requirements
- Deep technical expertise and original work in at least one relevant area.
- Experience with compiler engineering, optimization passes, IRs, lowering, code generation, LLVM, or MLIR.
- Experience with GPU programming or performance engineering using CUDA, HIP, Metal, Vulkan, or similar technologies.
- Knowledge of LLM inference engines, attention, MoE, parallelism, communication, or other systems-level LLM execution concepts.
- Experience with formal verification, equivalence checking, SAT/SMT, or systems that automatically generate, search, test, benchmark, or optimize code.
- Ability to explain technical decisions, alternatives, and measured results through code, papers, theses, projects, or detailed technical work.
Culture & Benefits
- Small engineering and research team of 10 people, including 9 engineers and researchers and 4 PhDs.
- Direct collaboration across the full inference stack.
- Fast iteration from optimization ideas to compilation, execution, verification, and measurement on real GPUs.
- High ownership over technical decisions and the systems shaping LLM inference optimization.
- Opportunity to deepen expertise in compilers, GPU systems, or LLM inference while expanding across the stack.
Hiring process
- Technical work is reviewed during the process, including public code, upstream contributions, papers, theses, technical projects, or detailed write-ups.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →