Назад
Company hidden
5 часов назад

Inference Engineer (AI)

195 000 - 285 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Inference Engineer (AI) (LLM inference and heterogeneous hardware): Building deployed, optimized inference systems, runtimes, serving frameworks, and evaluation tools for frontier AI models across D-Matrix silicon, CPUs, GPUs, and custom accelerators with an accent on kernel optimization, quantization, distributed inference, and performance engineering. Focus on designing proof-of-concept systems, managing KV caches and parallelism, and translating inference workloads into hardware and product improvements.

Location: Hybrid in Santa Clara, California, United States

Salary: $195,000–$285,000 annually at L6, plus equity, bonus, and benefits.

Company

hirify.global develops software and hardware for generative AI inference using its in-memory compute architecture and heterogeneous deployments.

What you will do

  • Identify and prototype emerging LLM inference use cases for heterogeneous hardware deployments.
  • Build proof-of-concept systems demonstrating hirify.global capabilities to customers, partners, and internal stakeholders.
  • Develop and tune custom kernels, operators, quantization, sparsity, and batching strategies to improve throughput and latency.
  • Build and maintain inference runtimes, serving frameworks, benchmarking, evaluation, and production tooling.
  • Contribute to distributed inference systems, including tensor and pipeline parallelism, disaggregated prefill/decode, and KV-cache management.
  • Collaborate with hardware architects, firmware and compiler teams, product, and business development to turn inference work into product demonstrations and roadmap insights.

Requirements

  • 10+ years of relevant engineering experience with a bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience; alternatively, 6+ years with a master's or PhD.
  • Strong proficiency in Python and C/C++.
  • Hands-on experience optimizing LLM inference, including attention kernels, KV cache, batching, and INT8, FP8, or INT4 quantization.
  • Contributor-level experience with an inference framework such as vLLM, SGLang, TensorRT-LLM, or ONNX Runtime.
  • Familiarity with GPU kernel programming using CUDA or Triton and performance profiling tools.
  • Ability to work in a hybrid role based in Santa Clara, California.

Nice to have

  • Experience with heterogeneous compute, custom silicon, or ASIC-based inference deployments.
  • Experience with distributed inference, production serving at scale, latency SLOs, continuous batching, or multi-model serving.
  • Knowledge of speculative decoding, mixture-of-experts routing, long-context serving, or systems-level LLM training and inference.
  • Contributions to open-source inference or machine learning systems projects.

Culture & Benefits

  • Work on novel hirify.global in-memory compute hardware and inference optimization problems.
  • End-to-end ownership from research ideas and prototypes to deployed systems.
  • Small, senior team with high autonomy and direct influence on product direction.
  • Medical, dental, vision, 401(k), equity, bonus, and additional wellbeing-focused benefits.
  • Inclusive environment emphasizing respect, collaboration, humility, direct communication, and execution.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →