13 дней назад
Member of Technical Staff — Model Optimization and Inference (New Grad)
200 000 - 300 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff — Model Optimization and Inference (New Grad) (AI): Optimizing inference for full-duplex multimodal AI avatar systems with an accent on sub-500ms latency, model serving, and memory efficiency. Focus on designing KV cache strategies, accelerating diffusion and LLM inference, applying quantization, and eliminating end-to-end performance bottlenecks.
Location: In-person in Seattle, Washington, five days a week
Salary: $200,000–$300,000 annual base salary, plus meaningful equity
Company
is a research company building photorealistic, real-time AI avatars with emotional intelligence through full-duplex audiovisual systems.
What you will do
- Optimize end-to-end inference across LLMs, audio models, and diffusion-based components.
- Implement KV cache eviction, compression, and memory-efficient attention for long-context conversations.
- Extend inference serving frameworks such as vLLM, SGLang, and TensorRT-LLM for multimodal real-time workloads.
- Profile and benchmark latency and throughput while identifying and eliminating bottlenecks.
- Build profiling viewers, inference test harnesses, and internal optimization infrastructure.
- Apply diffusion acceleration, custom kernels, quantization, and other techniques to improve throughput without materially reducing quality.
Requirements
- Completed or nearly completed BS, MS, or PhD in computer science, machine learning, or a related field.
- Strong fundamentals in LLM inference or ML systems, including KV caching, memory layout, attention kernels, batching, or serving.
- Exposure to vLLM, SGLang, TensorRT-LLM, or similar inference serving frameworks.
- Strong Python and PyTorch skills.
- A systematic approach to profiling and optimization, with a focus on measuring before optimizing.
- Ability to work in person in Seattle five days per week.
Nice to have
- CUDA or Triton experience and kernel optimization work.
- Experience with diffusion inference, speculative decoding, quantization, or model compression.
- Internship, research, publication, or open-source experience in LLM inference, ML systems, or model serving.
- Familiarity with multimodal or streaming inference architectures and hard latency SLAs.
Culture & Benefits
- Research-focused environment working on unsolved real-time AI systems problems.
- Visa sponsorship is available from day one, including O-1, H-1B, and green card sponsorship.
- Health Savings Account plan with approximately $2,000 in annual company contributions.
- 15 days of paid time off, public holidays, and a full week of office closure at year-end.
- Workday meals, drinks, snacks, commuter benefits, and a 401(k).
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
10 дней назад
Performance Engineer (AI)
210 000 - 250 000$
12 дней назад
Senior Research Engineer (AI)
270 000 - 310 000$
13 дней назад
Senior Staff Machine Learning Engineer (AI)
220 000 - 280 000$
Decagon
9 дней назад
Research Engineer (Audio and Speech)
200 000 - 400 000$
13 дней назад
Member of Technical Staff (Speech and Language Models)
200 000 - 300 000$
12 дней назад
Technical Lead, On-Device AI Inference
300 000 - 500 000$