обновлено 9 дней назад
Founding AI Inference Engineer (AI)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Founding AI Inference Engineer (AI): Define and build how Fuse serves AI inference workloads at scale with an accent on inference serving architecture, high-throughput/low-latency request handling, and reliability. Focus on translating performance commitments into serving capacity plans while integrating GPU/CUDA performance work into the serving layer.
Location: Remote; listed locations are London, England, United Kingdom and the United States
Company
is an energy startup building an integrated energy company spanning renewable generation, batteries, grid infrastructure, real-time power trading, distributed energy, and AI-powered high-performance compute infrastructure.
What you will do
- Define the inference serving strategy and architecture from first principles.
- Design and build request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive workloads.
- Own model-level optimization using quantization, distillation, speculative decoding, and related techniques.
- Make architecture decisions across serving frameworks and orchestration, including vLLM, TensorRT-LLM, SGLang, and Triton Inference Server.
- Translate throughput, latency, and uptime commitments into technical specifications and capacity plans.
- Own inference performance and reliability while integrating CUDA, GPU, custom kernel, and hardware work into the serving layer.
Requirements
- 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project or industry experience.
- Deep hands-on experience with inference serving frameworks, batching, KV-cache management, quantization, and speculative decoding.
- Strong systems thinking across the full path from incoming request to served response across a large cluster.
- Experience working directly with GPU and CUDA engineers on low-level performance integration.
- Experience making high-stakes architecture decisions and owning the outcomes.
- Comfort shaping a new function without an established playbook.
Nice to have
- Experience with Triton or custom ML inference and training frameworks.
- Autoscaling, large-scale inference capacity planning, multi-tenant serving, or SLA-driven infrastructure experience.
- Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
- Kubernetes or Slurm experience.
- Interest in energy markets, grid systems, or sustainability-focused compute.
Culture & Benefits
- Work in a founding role reporting directly to the CTO and shape a new engineering function.
- Competitive salary with equity eligibility and a biannual bonus scheme.
- Fully expensed technology matched to role requirements.
- Private health insurance.
- Benefits vary by location; breakfast and dinner allowances are available to office-based employees.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
9 дней назад
AI Infrastructure Engineer (AI)
11 дней назад
AI Engineer
Yandex
13 дней назад
Разработчик Inference Server на C++ в отдел ML-инфраструктуры (AI)
10 дней назад
Lead Machine Learning Engineer (AI)
9 дней назад
Engineering Manager, AI Models Infrastructure
12 дней назад