Назад
Company hidden
4 часа назад

Engineering Lead, Inference Optimization (AI)

270 000 - 330 000$
Формат работы
remote (только USA)
Тип работы
fulltime
Грейд
lead
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Engineering Lead, Inference Optimization (AI): Leading inference performance strategy and a small optimization team while building high-throughput, cost-efficient LLM inference infrastructure with an accent on GPU architecture, benchmarking, and distributed execution. Focus on optimizing latency, throughput, and cost per token, evaluating inference engines and hardware, and developing advanced quantization, parallelism, profiling, and kernel optimization techniques.

Location: Remote — US only

Base annual salary: $270,000–$330,000 USD

Company

hirify.global is a privacy-focused consumer AI company building a platform for individuals, applications, and AI agents.

What you will do

  • Own the technical strategy for inference performance.
  • Recruit and lead the Inference Optimization Team.
  • Optimize GPU infrastructure across architectures including H200 and B300 GPUs.
  • Improve latency, throughput, and cost per token for high-volume LLM inference workloads.
  • Build reproducible benchmarking harnesses for inference engines such as vLLM and SGLang, evaluating engines, quantization schemes, and parallelism strategies.
  • Optimize multivariate inference load balancing and evaluate emerging techniques, kernels, attention variants, compilation improvements, and inference hardware.

Requirements

  • 8+ years of experience in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge.
  • 5+ years of experience leading engineering teams, combining hands-on individual contribution with team management.
  • Proficiency in Python, Rust, or Go.
  • Production experience with at least one high-volume LLM inference engine, such as vLLM or SGLang.
  • Experience with continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile.
  • Experience with distributed inference, GPU profiling, and multi-GPU or multi-node environments.

Nice to have

  • C++ and CUDA experience.
  • Hands-on custom CUDA or Triton kernel development.
  • Diffusion or image model inference optimization experience.
  • Contributions to open-source inference frameworks.

Culture & Benefits

  • Work on privacy-focused AI with zero data retention and zero training on user inputs.
  • High-agency environment centered on curiosity, ownership, ethical principles, and collaboration.
  • Opportunity to shape technical strategy and build an inference optimization function from the ground up.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →