Назад
Company hidden
9 дней назад

AI Inference Engineer

250 000 - 300 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Inference Engineer (LLM Inference/GPU Systems): Building production-scale inference infrastructure for modern AI models with an accent on latency, throughput, cost efficiency, and distributed multi-GPU execution. Focus on optimizing batching, quantization, KV caching, parallelism, GPU scheduling, and Kubernetes-based workloads while solving compute, memory, and networking bottlenecks.

Location: San Francisco Bay Area, Santa Clara County, CA, Arlington, VA, Boulder, CO, Fremont, CA, Mountain View, CA, or New York

Salary: $250,000–$300,000 per year

Company

Stanford-spun AI company building a production AI platform.

What you will do

  • Architect high-performance infrastructure for serving LLMs at production scale.
  • Optimize inference latency, throughput, cost efficiency, batching, quantization, KV caching, and parallelism.
  • Build distributed execution across multi-GPU and multi-node environments.
  • Improve GPU scheduling, allocation, and utilization across Kubernetes-based infrastructure.
  • Profile compute, memory, and networking bottlenecks across the inference stack.
  • Make systems-level trade-offs and own meaningful parts of an inference stack built from the ground up.

Requirements

  • Strong background in ML systems, distributed computing, HPC, or performance engineering.
  • Strong Python and/or C++ skills.
  • Solid understanding of PyTorch and GPU computing.
  • Ability to reason about GPU utilization, memory bandwidth, communication overhead, and model architecture.
  • Experience with CUDA, NCCL, Triton, vLLM, TensorRT-LLM, or SGLang is highly relevant.

Culture & Benefits

  • Opportunity to build an inference platform from the ground up rather than maintain an established system.
  • Ownership of meaningful parts of the production AI infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →