Назад
Company hidden
7 часов Π½Π°Π·Π°Π΄

Developer Relations Engineer (AI Inference)

200Β 000 - 400Β 000$
Π€ΠΎΡ€ΠΌΠ°Ρ‚ Ρ€Π°Π±ΠΎΡ‚Ρ‹
remote (Ρ‚ΠΎΠ»ΡŒΠΊΠΎ USA)/onsite
Π’ΠΈΠΏ Ρ€Π°Π±ΠΎΡ‚Ρ‹
fulltime
Английский
b2
Π‘Ρ‚Ρ€Π°Π½Π°
US
Вакансия ΠΈΠ· списка Hirify.GlobalВакансия ΠΈΠ· Hirify Global, списка ΠΌΠ΅ΠΆΠ΄ΡƒΠ½Π°Ρ€ΠΎΠ΄Π½Ρ‹Ρ… tech-ΠΊΠΎΠΌΠΏΠ°Π½ΠΈΠΉ
Для мэтча ΠΈ ΠΎΡ‚ΠΊΠ»ΠΈΠΊΠ° Π½ΡƒΠΆΠ΅Π½ Plus

ΠœΡΡ‚Ρ‡ & Π‘ΠΎΠΏΡ€ΠΎΠ²ΠΎΠ΄

Для мэтча с этой вакансиСй Π½ΡƒΠΆΠ΅Π½ Plus

ОписаниС вакансии

ВСкст:
/
TL;DR
Developer Relations Engineer (AI Inference): Building technical deep dives, demos, tutorials, documentation, and workshops that help developers understand and scale vLLM-based AI inference with an accent on LLM inference systems, GPU serving, and distributed runtimes. Focus on explaining KV cache, batching, quantization, scheduling, and model-server tradeoffs while turning complex systems concepts into practical developer education.

Location: San Francisco, California; remote work may be considered within the US for exceptional candidates

Annual salary: $200,000–$400,000 USD plus equity

Company

hirify.global was founded by the creators and core maintainers of vLLM to make AI inference cheaper and faster and grow vLLM as a leading AI inference engine.

What you will do

  • Write technical deep dives, tutorials, documentation, examples, and architecture explainers for AI inference practitioners.
  • Build demos, benchmarks, runnable repositories, hands-on labs, and other public technical artifacts.
  • Teach concepts including KV cache, PagedAttention, continuous batching, prefix caching, prefill and decode scheduling, quantization, and speculative decoding.
  • Explain GPU serving, distributed runtimes, tensor and data parallelism, and latency-versus-throughput tradeoffs.
  • Host workshops and technical talks and help developers move from concepts to working code.
  • Contribute to developer-facing open-source education and shape adoption of vLLM across the AI infrastructure community.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, machine learning, systems, or a related field.
  • Strong technical understanding of LLM inference systems, model serving, GPU inference, distributed runtimes, scheduling, batching, quantization, or related infrastructure.
  • Experience with vLLM or adjacent technologies such as SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, or similar systems.
  • A strong public portfolio of technical artifacts, including blogs, tutorials, workshops, courses, open-source documentation, benchmark posts, conference talks, demos, or runnable repositories.
  • Ability to clearly explain complex systems concepts and create practical developer education without relying on generic content marketing.
  • Strong engineering judgment, product taste, and ability to turn raw technical material into useful learning resources.

Nice to have

  • Experience in ML systems, distributed systems, HPC, compilers, GPU kernels, serving infrastructure, MLOps, developer tooling, or open-source infrastructure.
  • Contributions to developer-facing open source through documentation, tutorials, examples, cookbooks, demos, or community support.
  • Community credibility in AI infrastructure, CUDA/GPU, Ray, vLLM, PyTorch, Modal, BentoML, Baseten, or related ecosystems.
  • Widely shared technical writing, courses, architecture deep dives, demos, benchmarks, or repositories focused on LLM inference and model serving.

Culture & Benefits

  • Research and Engineering environment focused on AI inference systems.
  • Visa sponsorship is available on a case-by-case basis.
  • Health, dental, and vision benefits.
  • 401(k) company match.
  • Equity included in the compensation package.

Π‘ΡƒΠ΄ΡŒΡ‚Π΅ остороТны: Ссли Ρ€Π°Π±ΠΎΡ‚ΠΎΠ΄Π°Ρ‚Π΅Π»ΡŒ просит Π²ΠΎΠΉΡ‚ΠΈ Π² ΠΈΡ… систСму, ΠΈΡΠΏΠΎΠ»ΡŒΠ·ΡƒΡ iCloud/Google, ΠΏΡ€ΠΈΡΠ»Π°Ρ‚ΡŒ ΠΊΠΎΠ΄/ΠΏΠ°Ρ€ΠΎΠ»ΡŒ, Π·Π°ΠΏΡƒΡΡ‚ΠΈΡ‚ΡŒ ΠΊΠΎΠ΄/ПО, Π½Π΅ Π΄Π΅Π»Π°ΠΉΡ‚Π΅ этого - это мошСнники. ΠžΠ±ΡΠ·Π°Ρ‚Π΅Π»ΡŒΠ½ΠΎ ΠΆΠΌΠΈΡ‚Π΅ "ΠŸΠΎΠΆΠ°Π»ΠΎΠ²Π°Ρ‚ΡŒΡΡ" ΠΈΠ»ΠΈ ΠΏΠΈΡˆΠΈΡ‚Π΅ Π² ΠΏΠΎΠ΄Π΄Π΅Ρ€ΠΆΠΊΡƒ. ΠŸΠΎΠ΄Ρ€ΠΎΠ±Π½Π΅Π΅ Π² Π³Π°ΠΉΠ΄Π΅ β†’