Назад
Company hidden
3 дня назад

Forward Deployed Engineer (AI)

200 000 - 400 000$
Формат работы
remote (только USA)/onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Forward Deployed Engineer (AI): Deploying, integrating, debugging, and optimizing vLLM-powered inference systems across customer cloud, Kubernetes, GPU, networking, and model-serving environments with an accent on production reliability and performance. Focus on solving cross-layer systems problems, optimizing latency and throughput, and turning recurring customer issues into reusable product capabilities.

Location: Based in San Francisco, California; remote work may be considered within the US for exceptional candidates.

Salary: $200,000–$400,000 annual salary plus equity.

Company

hirify.global develops vLLM-based AI inference infrastructure to make model inference faster and more cost-efficient.

What you will do

  • Deploy, integrate, debug, and optimize vLLM-powered inference systems in customer environments.
  • Implement solutions across cloud platforms, Kubernetes, GPUs, networking, and model-serving systems.
  • Own complex production issues from architecture discussions through implementation and resolution.
  • Work with customer engineering teams and core product and engineering teams.
  • Convert recurring customer problems into reusable product capabilities, tooling, and platform improvements.

Requirements

  • Bachelor’s degree or equivalent experience in computer science, engineering, systems, machine learning, or a related field.
  • Strong software engineering skills in Python, Go, TypeScript, or similar languages.
  • Production experience deploying or operating ML systems, model serving, AI infrastructure, cloud platforms, Kubernetes, or high-scale backend systems.
  • Strong debugging skills across applications, runtimes, infrastructure, networking, identity, storage, observability, and distributed systems.
  • Ability to reason about latency, throughput, batching, compatibility, scaling, reliability, and cost in production inference environments.
  • Ability to work directly with customer engineering teams and drive ambiguous technical implementations to resolution.

Nice to have

  • Experience with vLLM, SGLang, TensorRT-LLM, TGI, Ray Serve, BentoML, or similar inference systems.
  • Experience with NVIDIA or AMD GPUs, CUDA, ROCm, GPU scheduling, or multi-GPU serving.
  • Experience with enterprise, regulated, security-sensitive, or bring-your-own-cloud deployments.
  • Experience building APIs, SDKs, CLIs, deployment platforms, control planes, or infrastructure products.
  • Open-source contributions or experience in forward-deployed, customer engineering, field engineering, or technical solutions roles.

Culture & Benefits

  • Hands-on engineering role focused on implementation rather than traditional pre-sales.
  • Direct influence on customer time-to-value and product evolution.
  • Health, dental, and vision benefits.
  • 401(k) company match.
  • Visa sponsorship available on a case-by-case basis.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →