Назад
Company hidden
3 дня назад

AI Inference Intern

Формат работы
onsite
Тип работы
fulltime
Грейд
trainee
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Inference Intern (vLLM): Building and optimizing systems that make frontier AI model inference faster, cheaper, and more reliable across runtimes, accelerators, distributed serving, and cloud infrastructure with an accent on performance engineering, systems software, and measurable correctness. Focus on implementing inference techniques, optimizing GPU kernels, scaling multi-GPU and multi-node workloads, and contributing production-ready or open-source software.

Location: San Francisco, California; in-office only at hirify.global's San Francisco office

Compensation: Competitive compensation based on the applicable co-op market and candidate background, plus a housing stipend for the co-op term.

Company

hirify.global develops vLLM as an AI inference engine, focusing on making model inference faster, cheaper, and more reliable.

What you will do

  • Work with vLLM creators and maintainers on production, open-source, or supporting tooling.
  • Contribute to inference runtime features such as model architectures, scheduling, continuous batching, KV-cache management, and hybrid serving.
  • Build distributed serving systems for multi-GPU and multi-node workloads, including parallelism, fault tolerance, networking, and KV-cache transport.
  • Write and optimize GPU, accelerator, attention, GEMM, sampling, fused, and quantization kernels.
  • Develop backend, compiler, runtime, and performance integrations for AMD GPUs and Google TPUs.
  • Build cloud orchestration capabilities including Kubernetes operators, GPU scheduling, routing, observability, automated recovery, and multi-cloud infrastructure.

Requirements

  • Currently pursuing a bachelor's, master's, or PhD degree and eligible for a University of Waterloo co-op work term.
  • Strong programming ability in Python, C++, Rust, Go, or another systems-oriented language.
  • Strong computer science fundamentals and the ability to learn from research papers, technical documentation, and complex systems code.
  • Evidence of building and debugging nontrivial software through coursework, research, internships, open-source work, or ambitious projects.
  • Ability to define success, measure results, validate correctness, communicate clearly, and iterate on difficult systems problems.
  • Ability to work in person at the San Francisco office.

Nice to have

  • Depth in ML or inference systems, distributed systems, GPU or accelerator programming, compilers, high-performance computing, operating systems, networking, Kubernetes, or cloud infrastructure.
  • Experience with PyTorch, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, ROCm/HIP, JAX/XLA, Kubernetes, Helm, Terraform, Ray, or SLURM.
  • Experience with profiling, benchmarking, numerical or systems correctness, performance-regression testing, multi-GPU, multi-node, or production-like workloads.
  • Open-source contributions, research systems, benchmark suites, compiler or kernel projects, distributed services, or infrastructure tools.
  • Technical artifacts such as papers, design documents, demos, blog posts, or talks.

Culture & Benefits

  • Mentorship from a small, senior team of vLLM creators and core maintainers.
  • Meaningful ownership and opportunities to collaborate across models, compilers, accelerators, networking, and distributed systems.
  • Work intended for open source, production systems, or supporting infrastructure.
  • Competitive co-op compensation and a housing stipend for the co-op term.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →