Назад
Company hidden
6 дней назад

LLM Inference Engineer

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
LLM Inference Engineer (AI/ML): Making large language models run faster, cheaper, and more reliably in production with an accent on inference optimization, serving architectures, and customer-specific deployments. Focus on profiling vLLM and SGLang down to CUDA kernels, tuning latency, throughput, and cost, and taking workloads from proof of concept to monitored production services.

Location: San Francisco, CA, United States; on-site

Compensation: up to $250,000–$300,000 plus bonus and Restricted Stock Units.

Company

hirify.global builds vertically integrated AI infrastructure across energy, data centers, and cloud services to support demanding AI workloads.

What you will do

  • Bring modern large language model inference techniques into production and refine them for real customer workloads.
  • Design and optimize serving architectures, including prefill and decode disaggregation and request routing.
  • Profile and improve performance across serving frameworks such as vLLM and SGLang and the CUDA kernels underneath.
  • Adapt optimization methods across ML models and tune deployments for latency, throughput, cost, and reliability.
  • Partner with customer engineering teams to move workloads from proof of concept to monitored production services.
  • Build, test, and support inference software and product features from experimentation through production delivery.

Requirements

  • Bachelor’s, master’s, or Ph.D. in computer science, engineering, mathematics, or a related field.
  • Production software engineering experience with Python, C++, or another general-purpose language; Python is preferred.
  • Hands-on experience optimizing large language models for high-throughput and low-latency inference.
  • Experience with vLLM or SGLang, performance profiling, and kernel-level analysis.
  • Strong understanding of GPU architecture and behavior, AI/ML pipelines, and model development and deployment.
  • Strong communication skills for explaining complex technical topics to customers and teammates.

Nice to have

  • Experience with CUDA or comparable technologies.
  • Track record of improving software performance, particularly for large language models.
  • Experience building AI/ML inference systems and working with Docker and Kubernetes.
  • Customer-facing experience tuning AI/ML projects.

Culture & Benefits

  • Competitive compensation, bonus eligibility, equity, and Restricted Stock Units.
  • Paid time off, holidays, leave programs, and paid parental leave.
  • Health, dental, vision, HSA contributions, life insurance, and disability coverage.
  • Professional development, tuition reimbursement, and mental health and wellness support.
  • 401(k) plan with company matching up to 4% of salary.
  • Commuter benefits, daily meal allowance, cell phone stipend, volunteer time off, and global travel insurance.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →