Назад
Company hidden
9 часов назад

Generalist Engineer (AI)

Формат работы
remote (Global)
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Generalist Engineer (AI) (vLLM/inference infrastructure): Building and optimizing the full AI inference stack, from GPU kernels and model execution to distributed systems and cloud orchestration, with an accent on performance, scale, and autonomous end-to-end delivery. Focus on designing low-level CUDA or equivalent kernels, serving models across thousands of accelerators, and building reliable Kubernetes-based infrastructure for global AI inference.

Location: Fully remote, worldwide; timezone-flexible with regular overlap with Pacific Time for critical syncs.

Company

hirify.global develops and advances vLLM as an AI inference engine, focusing on making model inference faster and more cost-efficient.

What you will do

  • Optimize LLM and diffusion model serving within the vLLM inference runtime.
  • Develop CUDA, Triton, TileLang, Pallas, or equivalent kernels for diverse accelerator architectures.
  • Build distributed systems that serve models across thousands of accelerators with minimal latency.
  • Develop cloud orchestration, cluster management, deployment automation, and production monitoring infrastructure.
  • Work across the vLLM stack, including GPU kernels, distributed systems, model architectures, and ML infrastructure.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, or a related field.
  • Demonstrated autonomy and ability to drive complex projects from vague problem statements to shipped code.
  • Strong asynchronous communication skills and ability to collaborate across time zones.
  • Strong record of delivering high-impact work in complex technical environments.
  • Deep expertise in systems programming, GPU or accelerator programming, distributed systems, or ML infrastructure.
  • Strong proficiency in at least two areas including GPU architecture and kernels, Rust/Go/C++ distributed systems, Python with PyTorch and LLM inference, Kubernetes, or transformer-based model serving.

Nice to have

  • Contributions to vLLM or major open-source ML or systems projects.
  • Experience with NVIDIA, AMD, TPU, or Intel accelerators.
  • Knowledge of quantization, ML kernel optimization, or compiler technologies.
  • Experience improving reliability and performance at scale.
  • Technical writing or impactful ML infrastructure side projects.

Culture & Benefits

  • Asynchronous, autonomous work with regular Pacific Time overlap for critical coordination.
  • Collaboration with the San Francisco headquarters in a globally remote environment.
  • Competitive compensation comprising salary and equity, adjusted to local market conditions.
  • Location-appropriate benefits, including health coverage where applicable.
  • Visa sponsorship available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →