Назад
Company hidden
обновлено 9 дней назад

Founding AI Inference Engineer (AI)

Формат работы
remote (Global)
Тип работы
fulltime
Английский
b2
Страна
UK/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Founding AI Inference Engineer (AI): Define and build how Fuse serves AI inference workloads at scale with an accent on inference serving architecture, high-throughput/low-latency request handling, and reliability. Focus on translating performance commitments into serving capacity plans while integrating GPU/CUDA performance work into the serving layer.

Location: Remote; listed locations are London, England, United Kingdom and the United States

Company

hirify.global is an energy startup building an integrated energy company spanning renewable generation, batteries, grid infrastructure, real-time power trading, distributed energy, and AI-powered high-performance compute infrastructure.

What you will do

  • Define the inference serving strategy and architecture from first principles.
  • Design and build request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive workloads.
  • Own model-level optimization using quantization, distillation, speculative decoding, and related techniques.
  • Make architecture decisions across serving frameworks and orchestration, including vLLM, TensorRT-LLM, SGLang, and Triton Inference Server.
  • Translate throughput, latency, and uptime commitments into technical specifications and capacity plans.
  • Own inference performance and reliability while integrating CUDA, GPU, custom kernel, and hardware work into the serving layer.

Requirements

  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project or industry experience.
  • Deep hands-on experience with inference serving frameworks, batching, KV-cache management, quantization, and speculative decoding.
  • Strong systems thinking across the full path from incoming request to served response across a large cluster.
  • Experience working directly with GPU and CUDA engineers on low-level performance integration.
  • Experience making high-stakes architecture decisions and owning the outcomes.
  • Comfort shaping a new function without an established playbook.

Nice to have

  • Experience with Triton or custom ML inference and training frameworks.
  • Autoscaling, large-scale inference capacity planning, multi-tenant serving, or SLA-driven infrastructure experience.
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
  • Kubernetes or Slurm experience.
  • Interest in energy markets, grid systems, or sustainability-focused compute.

Culture & Benefits

  • Work in a founding role reporting directly to the CTO and shape a new engineering function.
  • Competitive salary with equity eligibility and a biannual bonus scheme.
  • Fully expensed technology matched to role requirements.
  • Private health insurance.
  • Benefits vary by location; breakfast and dinner allowances are available to office-based employees.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →