Назад
Company hidden
4 дня назад

Performance Engineer (AI)

210 000 - 250 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Performance Engineer (AI): Building and maintaining reproducible performance and energy benchmarks for an optical AI inference accelerator with an accent on workload fidelity, GPU measurement, and energy efficiency. Focus on correlating roofline models, architecture models, RTL simulations, and measured competitor hardware while analyzing LLM inference latency, throughput, and power.

Location: Full-time onsite in Austin, Texas or Sunnyvale, California

Salary: $210K–$250K plus equity

Company

hirify.global is an AI hardware startup developing silicon photonics and programmable metasurface technology for energy-efficient optical inference accelerators.

What you will do

  • Own the performance and energy metrics used for architecture, product, and leadership decisions.
  • Produce consistent measurements across roofline and limiter models, architecture performance models, RTL simulation, and competitor hardware.
  • Define and maintain workload methodology across model, sequence length, batch, precision, inference phase, and parallelism.
  • Bring up inference workloads from Hugging Face, PyTorch, research papers, and serving stacks including vLLM, SGLang, TensorRT-LLM, and Triton Inference Server.
  • Measure competing GPUs and accelerators end to end, including cloud or lab environments, drivers, images, and run recipes.
  • Report TTFT, inter-token latency, throughput, throughput per watt, and energy, with reproducible harnesses, configurations, plots, and logs.

Requirements

  • BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
  • 5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement.
  • Experience building or operating benchmark harnesses that generate measured results on real GPUs or accelerators.
  • Hands-on experience with roofline analysis, limiter analysis, analytical performance modeling, and GPU profiling using NVIDIA Nsight Systems, Nsight Compute, or equivalent tools.
  • Working knowledge of LLM inference stacks, including prefill versus decode, continuous batching, and Mixture-of-Experts models.
  • Proficiency in Python and Linux, plus experience with cloud GPU operations, containers, instance types, drivers, quotas, and cost management.

Nice to have

  • Experience correlating performance models or RTL/Verilator simulations with measured silicon or GPUs.
  • GPU kernel experience with CUDA, CUTLASS, or Triton, or familiarity with PyTorch internals.
  • Knowledge of PagedAttention, FlashAttention, speculative decoding, and disaggregated prefill.
  • Distributed inference experience with collectives, all-reduce, NCCL, NVLink, InfiniBand, MLPerf, or production benchmarking pipelines.
  • Background at a hyperscaler, GPU vendor, accelerator company, or inference lab.

Culture & Benefits

  • Collaborative engineering environment focused on photonics, AI, and computational performance.
  • 100% coverage of base health plan premiums for employees and dependents, plus HSA contributions.
  • Unlimited PTO.
  • 401(k) matching and stock option opportunities.
  • Dental, vision, life, hospital, critical illness, and accident insurance, with personalized plan options.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →