Назад
Company hidden
17 часов назад

Member of Technical Staff — Model Optimization and Inference (Experienced)

250 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — Model Optimization and Inference (Experienced) (AI inference): Optimizing real-time inference across LLM, audio, and diffusion model stacks with an accent on sub-500ms latency, throughput, KV caching, quantization, and serving infrastructure. Focus on profiling bottlenecks, accelerating multimodal full-duplex systems, and developing efficient kernels and deployment strategies for production workloads.

Location: In-person in Seattle, Washington, five days a week

Salary: $250,000–$350,000 base salary per year, plus equity

Company

hirify.global is a research company building photorealistic, real-time AI avatars with emotional intelligence through full-duplex audiovisual foundation models.

What you will do

  • Own end-to-end inference optimization across LLM, audio, and diffusion model stacks.
  • Implement KV cache strategies for long-context conversations, including eviction, compression, and memory-efficient attention.
  • Evaluate, deploy, and extend vLLM, SGLang, TensorRT-LLM, and similar inference serving frameworks.
  • Profile and benchmark latency and throughput, systematically identifying and eliminating bottlenecks.
  • Build profiling viewers, inference test harnesses, and internal tooling for rigorous optimization.
  • Accelerate diffusion inference and apply quantization, custom kernels, batching, and other inference-time optimization techniques.

Requirements

  • Significant hands-on experience optimizing LLM inference in production or high-traffic research environments.
  • Proficiency with vLLM, SGLang, TensorRT-LLM, or similar frameworks, including adapting non-default configurations to specialized workloads.
  • Experience optimizing diffusion model inference through latency reduction, step distillation, caching, or kernel-level work.
  • Strong Python and PyTorch skills, with a systematic measure-first approach to profiling and optimization.
  • Experience with KV caching, memory layout, attention kernels, batching strategies, or speculative decoding.
  • Visa sponsorship is available from day one, including O-1, H-1B, and green card sponsorship.

Nice to have

  • Experience with post-training quantization such as GPTQ or AWQ.
  • Familiarity with multimodal or streaming inference architectures.
  • Experience deploying real-time AI systems with hard latency SLAs.
  • Experience at an AI lab, inference startup, or high-traffic model serving platform.
  • Open-source contributions to inference frameworks.

Culture & Benefits

  • Small research-focused team working on unsolved real-time AI problems.
  • AI-native tooling and access to advanced development tools.
  • HSA health plan with approximately $2,000 in annual company contributions.
  • 15 days of PTO, public holidays, and a full week of office closure at year-end.
  • Weekday meals, drinks, snacks, commuter benefits, and a 401(k) plan.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →