Назад
Company hidden
36 минут назад

Senior Performance Engineer (AI Inference)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Performance Engineer (AI Inference): Building reproducible benchmarks and competitive pricing models for Cerebras inference workloads with an accent on latency, throughput, transformer architecture, and GPU optimization. Focus on designing fair workload comparisons, evaluating kernel and quantization techniques, and translating performance and total-cost-of-ownership data into enterprise pricing recommendations.

Location: Headquarters/Sunnyvale Office; on-site

Company

hirify.global builds large-scale AI hardware and inference systems designed to deliver significantly higher training and inference speeds than GPU-based platforms.

What you will do

  • Build, run, and maintain reproducible inference benchmarks for code generation, summarization, multi-turn conversation, and agentic tool use.
  • Measure tokens per second, time to first token, latency under concurrency, and total cost of ownership for real customer workloads.
  • Evaluate inference optimization techniques including kernel fusions, flash-attention variants, quantization, and GPU memory strategies.
  • Build and continuously update competitive pricing models covering token-based, throughput-based, and enterprise contract pricing.
  • Prepare competitive analyses and actionable briefs for Sales, Product, Engineering, executives, and sales leaders.
  • Track third-party benchmarking sources and ensure hirify.global performance is accurately represented.

Requirements

  • Deep practical experience with open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.
  • 5+ years of experience in ML systems, ML research engineering, or high-performance computing.
  • Strong understanding of LLM inference economics, including tokens, throughput, latency, batch sizes, precision trade-offs, and customer cost.
  • Strong understanding of transformer architecture internals, including attention mechanisms and KV-cache management.
  • Self-directed and resourceful approach to technical research and analysis.

Nice to have

  • ML research background with publications or significant open-source contributions focused on systems or efficiency.
  • Contributions to open-source inference or kernel optimization projects.
  • Excellent communication skills for collaborating with executives, engineers, and sales leaders.

Culture & Benefits

  • Opportunity to build an AI platform beyond the constraints of GPUs.
  • Access to cutting-edge AI research and opportunities to publish or open source work.
  • Work with one of the fastest AI supercomputers in the world.
  • Job stability combined with startup vitality.
  • Non-corporate work culture focused on individual beliefs, learning, growth, and support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →