4 часа назад
Member of Technical Staff — Developer Technology (AI)
200 000 - 400 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Member of Technical Staff — Developer Technology (AI): Accelerating LLM inference and training on modern GPU hardware through SGLang and Miles, with an accent on GPU profiling, custom kernels, model enablement, and distributed systems. Focus on root-causing performance bottlenecks, optimizing CUDA/ROCm/Triton workloads, supporting new models and silicon, and translating complex partner problems into reproducible technical solutions.
Location: Palo Alto, CA, United States
Salary: $200,000–$400,000 USD annually, plus equity
Company
is an infrastructure-first AI company building open systems for high-performance LLM inference and training, including SGLang and Miles.
What you will do
- Profile and optimize GPU performance for production AI workloads across current and next-generation hardware.
- Root-cause performance and correctness bottlenecks from GPU kernels through distributed multi-node systems.
- Specialize in inference performance, GPU kernels and model enablement, speculative decoding, or training systems.
- Build and optimize custom CUDA, ROCm, or Triton kernels and support new models on new hardware.
- Partner with expert engineering teams to turn ambiguous technical problems into measurable improvements, cookbooks, and recommendations.
- Feed user-driven improvements into the open-source SGLang and Miles systems and their future roadmap.
Requirements
- 4+ years of experience in GPU systems, LLM infrastructure, or performance engineering.
- Strong profiling, debugging, and performance root-cause analysis skills.
- Hands-on experience with at least one of CUDA, ROCm, or Triton, with willingness to work across platforms.
- Strong Python programming skills plus C++ or CUDA experience.
- Ability to make progress on difficult, ambiguous problems and ramp quickly into unfamiliar systems and codebases.
- Ability to translate ambiguous requests into clear technical plans, verified cookbooks, and actionable recommendations for expert engineering audiences.
Nice to have
- Experience with LLM inference internals, including distributed serving, parallelism, routing, KV-cache management, and scheduling.
- Experience with low-precision quantization such as FP8, INT8, or INT4, and custom GPU kernel optimization.
- Familiarity with speculative decoding, large-scale distributed training, elasticity, or long-context workloads.
- Experience optimizing across NVIDIA and AMD platforms.
- Experience with SGLang, Miles, vLLM, TensorRT-LLM, Megatron, or comparable frameworks, including open-source AI/ML contributions.
Culture & Benefits
- Work on open AI infrastructure intended to make frontier-level inference and training more accessible.
- Collaborate with leading AI companies, research labs, and expert engineering partners.
- Contribute to systems serving large-scale production workloads and distributed training environments.
- Equity is included in the compensation package.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
4 часа назад
Member of Technical Staff, Kernels (AI)
200 000 - 350 000$
3 часа назад
Member of Technical Staff, Infrastructure & Training Systems (AI)
5 часов назад
Software Engineer (AI)
175 000 - 220 000$
6 дней назад
Helix AI Engineer, Training Performance (CUDA)
200 000 - 400 000$
3 часа назад
Member of Technical Staff, Inference (AI)
5 часов назад
Member of Technical Staff (Research Infrastructure)
200 000 - 400 000$