Назад
Company hidden
6 часов назад

Performance Engineer (Inference, Training & GPU)

200 000 - 300 000$
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Performance Engineer (Inference, Training & GPU) (AI/GPU systems): Optimizing foundational 3D world models for high-throughput inference, production serving, and efficient training with an accent on CUDA and Triton kernels, low-precision execution, and GPU utilization. Focus on profiling bottlenecks, designing performance models, improving distributed training and serving efficiency, and preserving numerical correctness across hardware and precision changes.

Location: San Francisco

Salary: $200,000–$300,000 base salary annually, plus equity awards

Company

hirify.global builds foundational world models that perceive, generate, reason about, and interact with the 3D world through spatial intelligence.

What you will do

  • Optimize end-to-end inference and serving for latency, throughput, batching, caching, and scheduling at production scale.
  • Write and tune CUDA and Triton GPU kernels, including kernel fusion, memory and bandwidth optimization, and FP8/INT8 execution.
  • Improve training throughput and GPU utilization through parallelism, communication and compute overlap, mixed precision, and pipeline-stall reduction.
  • Build performance models, profiling workflows, and observability for throughput, latency, cost, utilization, and trade-offs.
  • Maintain numerical correctness across precision, kernel, and hardware changes.
  • Partner with researchers to productionize models and accelerate experiments.

Requirements

  • Strong foundations in profiling, roofline analysis, latency and throughput optimization, and root-cause investigation.
  • Deep GPU programming and optimization experience with CUDA and/or Triton.
  • Hands-on experience optimizing large-model inference and serving, including batching, KV/prompt caching, quantization, and low-latency sampling.
  • Hands-on experience with training performance, parallelism, distributed communication, mixed or low precision, and GPU utilization.
  • Working knowledge of PyTorch and/or JAX internals, including torch.compile, XLA, or similar compiler paths.
  • Strong Python proficiency with the ability to work in C++/CUDA and, as needed, Rust or Go.

Nice to have

  • Experience at an AI lab or ML-native company and productionizing research systems.
  • Experience with FP8/INT8 quantization, mixed precision, and numerical regression detection.
  • Distributed training and inference systems, including NCCL, NVLink, model and tensor parallelism, and fault tolerance.
  • Experience with generative, diffusion, video, or 3D/spatial models.
  • Multi-accelerator experience with GPUs, TPUs, or Trainium, and performance-modeling or observability frameworks.

Culture & Benefits

  • Hands-on individual-contributor role with direct ownership of shipped performance improvements.
  • Collaborative work across AI research, systems engineering, and product design.
  • Mission focused on spatially intelligent AI systems and 3D world models.
  • Base salary plus equity awards.
  • Equal employment opportunity and reasonable accommodations throughout the hiring process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →