Назад
Company hidden
13 дней назад

Member of Technical Staff — Model Optimization and Inference (New Grad)

200 000 - 300 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
junior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — Model Optimization and Inference (New Grad) (AI): Optimizing inference for full-duplex multimodal AI avatar systems with an accent on sub-500ms latency, model serving, and memory efficiency. Focus on designing KV cache strategies, accelerating diffusion and LLM inference, applying quantization, and eliminating end-to-end performance bottlenecks.

Location: In-person in Seattle, Washington, five days a week

Salary: $200,000–$300,000 annual base salary, plus meaningful equity

Company

hirify.global is a research company building photorealistic, real-time AI avatars with emotional intelligence through full-duplex audiovisual systems.

What you will do

  • Optimize end-to-end inference across LLMs, audio models, and diffusion-based components.
  • Implement KV cache eviction, compression, and memory-efficient attention for long-context conversations.
  • Extend inference serving frameworks such as vLLM, SGLang, and TensorRT-LLM for multimodal real-time workloads.
  • Profile and benchmark latency and throughput while identifying and eliminating bottlenecks.
  • Build profiling viewers, inference test harnesses, and internal optimization infrastructure.
  • Apply diffusion acceleration, custom kernels, quantization, and other techniques to improve throughput without materially reducing quality.

Requirements

  • Completed or nearly completed BS, MS, or PhD in computer science, machine learning, or a related field.
  • Strong fundamentals in LLM inference or ML systems, including KV caching, memory layout, attention kernels, batching, or serving.
  • Exposure to vLLM, SGLang, TensorRT-LLM, or similar inference serving frameworks.
  • Strong Python and PyTorch skills.
  • A systematic approach to profiling and optimization, with a focus on measuring before optimizing.
  • Ability to work in person in Seattle five days per week.

Nice to have

  • CUDA or Triton experience and kernel optimization work.
  • Experience with diffusion inference, speculative decoding, quantization, or model compression.
  • Internship, research, publication, or open-source experience in LLM inference, ML systems, or model serving.
  • Familiarity with multimodal or streaming inference architectures and hard latency SLAs.

Culture & Benefits

  • Research-focused environment working on unsolved real-time AI systems problems.
  • Visa sponsorship is available from day one, including O-1, H-1B, and green card sponsorship.
  • Health Savings Account plan with approximately $2,000 in annual company contributions.
  • 15 days of paid time off, public holidays, and a full week of office closure at year-end.
  • Workday meals, drinks, snacks, commuter benefits, and a 401(k).

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →