Назад
Company hidden
4 часа назад

Member of Technical Staff — Inference-Core Engine (AI)

200 000 - 400 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — Inference-Core Engine (AI): Designing and building large-scale inference systems for frontier AI models with an accent on latency, throughput, GPU utilization, and production reliability. Focus on optimizing model-serving runtimes, batching and scheduling strategies, memory management, and performance bottlenecks across the model, kernel, runtime, and network layers.

Location: Palo Alto, California, United States

Annual salary: $200,000–$400,000 USD plus equity

Company

hirify.global is an infrastructure-first AI company building open systems for frontier-model inference and training.

What you will do

  • Design and build large-scale inference systems for frontier AI models.
  • Optimize latency, throughput, GPU utilization, and cost in production inference.
  • Develop model-serving architectures, runtimes, batching, scheduling, and memory-management strategies.
  • Collaborate with kernel, compiler, and systems teams on performance optimization.
  • Debug bottlenecks across model, runtime, kernel, network, and system layers.
  • Build observability, profiling, and performance-analysis tooling while driving infrastructure reliability and scalability.

Requirements

  • 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems.
  • Strong expertise in large-scale inference systems for LLMs or generative models.
  • Deep understanding of GPU architecture, distributed systems, networking, and compute-intensive workload optimization.
  • Experience optimizing latency- and throughput-critical production systems and debugging across system layers.
  • Proficiency in Python, Rust, C++, or Go for production systems.

Nice to have

  • Experience with SGLang, vLLM, TensorRT-LLM, CUDA, Triton, or custom kernel optimization.
  • Open-source contributions in ML or systems infrastructure.
  • Experience with batching, KV-cache management, scheduling strategies, inference at 1,000+ GPUs, HPC, or high-performance systems.

Culture & Benefits

  • Infrastructure-first environment led by engineers with experience at xAI and NVIDIA.
  • Work on open systems supporting frontier-level AI infrastructure.
  • Equity included in the compensation package.
  • Equal employment opportunity regardless of protected personal characteristics.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →