Developer Relations Engineer (AI Inference)
ΠΡΡΡ & Π‘ΠΎΠΏΡΠΎΠ²ΠΎΠ΄
ΠΠ»Ρ ΠΌΡΡΡΠ° Ρ ΡΡΠΎΠΉ Π²Π°ΠΊΠ°Π½ΡΠΈΠ΅ΠΉ Π½ΡΠΆΠ΅Π½ Plus
ΠΠΏΠΈΡΠ°Π½ΠΈΠ΅ Π²Π°ΠΊΠ°Π½ΡΠΈΠΈ
Location: San Francisco, California; remote work may be considered within the US for exceptional candidates
Annual salary: $200,000β$400,000 USD plus equity
Company
was founded by the creators and core maintainers of vLLM to make AI inference cheaper and faster and grow vLLM as a leading AI inference engine.
What you will do
- Write technical deep dives, tutorials, documentation, examples, and architecture explainers for AI inference practitioners.
- Build demos, benchmarks, runnable repositories, hands-on labs, and other public technical artifacts.
- Teach concepts including KV cache, PagedAttention, continuous batching, prefix caching, prefill and decode scheduling, quantization, and speculative decoding.
- Explain GPU serving, distributed runtimes, tensor and data parallelism, and latency-versus-throughput tradeoffs.
- Host workshops and technical talks and help developers move from concepts to working code.
- Contribute to developer-facing open-source education and shape adoption of vLLM across the AI infrastructure community.
Requirements
- Bachelorβs degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
- Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.
- Experience with vLLM or adjacent technologies such as SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, or similar systems.
- A strong public portfolio of technical artifacts, including blogs, tutorials, workshops, courses, open-source documentation, benchmark posts, conference talks, demos, or runnable repositories.
- Ability to clearly explain complex systems concepts and create practical developer education without relying on generic content marketing.
- Strong engineering judgment, product taste, and ability to turn raw technical material into useful learning resources.
Nice to have
- Experience in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.
- Contributions to developer-facing open source through documentation, tutorials, examples, cookbooks, demos, or community support.
- Community credibility in AI infrastructure, CUDA/GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, or related ecosystems.
- Widely shared technical writing, courses, architecture deep dives, demos, benchmarks, or repositories focused on LLM inference and model serving.
Culture & Benefits
- Research and Engineering environment focused on AI inference systems.
- Visa sponsorship is available on a case-by-case basis.
- Health, dental, and vision benefits.
- 401(k) company match.
- Equity included in the compensation package.
ΠΡΠ΄ΡΡΠ΅ ΠΎΡΡΠΎΡΠΎΠΆΠ½Ρ: Π΅ΡΠ»ΠΈ ΡΠ°Π±ΠΎΡΠΎΠ΄Π°ΡΠ΅Π»Ρ ΠΏΡΠΎΡΠΈΡ Π²ΠΎΠΉΡΠΈ Π² ΠΈΡ ΡΠΈΡΡΠ΅ΠΌΡ, ΠΈΡΠΏΠΎΠ»ΡΠ·ΡΡ iCloud/Google, ΠΏΡΠΈΡΠ»Π°ΡΡ ΠΊΠΎΠ΄/ΠΏΠ°ΡΠΎΠ»Ρ, Π·Π°ΠΏΡΡΡΠΈΡΡ ΠΊΠΎΠ΄/ΠΠ, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡΠ΅ ΡΡΠΎΠ³ΠΎ - ΡΡΠΎ ΠΌΠΎΡΠ΅Π½Π½ΠΈΠΊΠΈ. ΠΠ±ΡΠ·Π°ΡΠ΅Π»ΡΠ½ΠΎ ΠΆΠΌΠΈΡΠ΅ "ΠΠΎΠΆΠ°Π»ΠΎΠ²Π°ΡΡΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡΠΈΡΠ΅ Π² ΠΏΠΎΠ΄Π΄Π΅ΡΠΆΠΊΡ. ΠΠΎΠ΄ΡΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β