Назад
Company hidden
1 день назад

Forward Deployed Engineer - AI Inference

250 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Forward Deployed Engineer - AI Inference (LLM inference/GPU infrastructure): Building and deploying high-performance inference workloads for open models with an accent on profiling, benchmarking, optimization, and production reliability. Focus on identifying bottlenecks across models, serving engines, kernels, accelerators, and infrastructure while owning customer performance outcomes.

Location: San Francisco, onsite

Salary: Up to $250,000 base salary plus equity.

Company

A well-funded AI infrastructure startup building a high-performance inference cloud for open models.

What you will do

  • Profile, benchmark, and trace customer inference workloads.
  • Identify bottlenecks across models, serving engines, kernels, hardware, and production infrastructure.
  • Prove performance using customer models and production traffic.
  • Build and deploy the engineering needed to bring workloads into production.
  • Own latency, throughput, reliability, and error-rate outcomes across three to six customer accounts.
  • Convert recurring customer problems into improvements for the core platform.

Requirements

  • Strong systems engineering background with the ability to work directly with technical customers.
  • Understanding of LLM inference fundamentals, including prefill, decode, latency, throughput, and related trade-offs.
  • Experience with profiling, tracing, benchmarking, and identifying production bottlenecks.
  • Ability to use GPU programming technologies such as CUDA, HIP, or Triton.
  • Ability to explain technical constraints clearly, defend engineering methodology, and adapt conclusions based on evidence.
  • Comfort working autonomously across the stack and taking responsibility for production outcomes.

Nice to have

  • Experience with LLM inference and model serving.
  • Distributed systems and infrastructure experience.
  • Performance engineering and optimization experience.
  • Knowledge of quantization, speculative decoding, heterogeneous accelerators, cluster operations, reliability, and observability.

Culture & Benefits

  • Highly autonomous and technically intensive environment.
  • Small, talent-dense team with direct influence on customer performance, product direction, and company growth.
  • Meaningful equity included with the compensation package.
  • Engineers move across the stack and solve problems without waiting for tightly defined specifications.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →