Назад
5 часов назад

Software Engineer (LLM Inference)

Формат работы
remote (только Europe)/onsite/hybrid
Тип работы
fulltime
Грейд
senior/lead/head
Английский
b2
Страна
Slovenia
Релокация
Slovenia
Вакансия от Hirify. Размещена напрямую Вакансия размещена на Hirify напрямую от HR/нанимающего менеджера

Мэтч

Покажет вашу совместимость с вакансией

Описание вакансии

TL;DR
Software Engineer (LLM Inference): Building and optimizing high-performance inference stacks for large language models with an accent on low-latency execution and GPU utilization. Focus on implementing CUDA/Triton kernels, managing KV-cache, and scaling distributed inference for production-grade AI systems.

About the role

Soniox is pushing the boundaries of real-time AI, and we’re looking for an engineer to help us run large language models with exceptional speed, efficiency, and reliability at production scale.

In this role, you’ll work deep in the LLM inference stack, from vLLM scheduling and KV-cache management to CUDA kernels, distributed execution, and GPU profiling, optimizing every part of the path from request to generated token.

In this role, you will:

  • Build and optimize our vLLM-based inference stack for low latency, high throughput, and maximum GPU utilization.
  • Optimize continuous batching, scheduling, prefill/decode, prefix caching, and KV-cache allocation and reuse.
  • Optimize or implement CUDA and Triton kernels for attention, GEMMs, sampling, normalization, and other critical model operations.
  • Evaluate and integrate technologies such as FlashAttention, FlashInfer, CUDA Graphs, torch.compile, speculative decoding, and quantization.
  • Optimize distributed inference using tensor, data, and expert parallelism, NCCL, NVLink/NVSwitch, and InfiniBand.
  • Work closely with researchers to bring new dense and MoE model architectures into production quickly and efficiently.

You might thrive in this role if you:

  • Have deep hands-on experience with LLM inference, ideally working inside vLLM, SGLang, TensorRT-LLM, or similar systems, not just deploying them.
  • Understand TTFT, inter-token latency, throughput, continuous batching, PagedAttention, KV caching, and prefill vs. decode performance.
  • Are comfortable profiling GPUs and reasoning about compute, memory bandwidth, kernel launches, synchronization, and communication bottlenecks.
  • Have experience with CUDA, Triton, PyTorch, NCCL, and modern NVIDIA GPU architectures.
  • Understand Transformer internals including MHA/GQA, RoPE, KV cache, quantization, and MoE.
  • Have experience optimizing distributed, performance-critical systems in production.
  • Care deeply about performance, simplicity, and reliability, and take ownership from profiling through production deployment.

Why Soniox

  • You’ll help build one of the most technically advanced voice AI platforms in the world, and push LLM inference performance at every layer of the stack.
  • You’ll work directly with a world-class team of engineers and researchers on hard, measurable problems spanning models, GPU kernels, distributed systems, and production infrastructure.
  • You'll have a voice in how our technology evolves, how our company grows, and how AI transforms human communication.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →

Вакансия размещена на Hirify напрямую от HR/нанимающего менеджера