Назад
Company hidden
4 часа назад

Member of Technical Staff — Developer Technology (AI)

200 000 - 400 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — Developer Technology (AI): Accelerating LLM inference and training on modern GPU hardware through SGLang and Miles, with an accent on GPU profiling, custom kernels, model enablement, and distributed systems. Focus on root-causing performance bottlenecks, optimizing CUDA/ROCm/Triton workloads, supporting new models and silicon, and translating complex partner problems into reproducible technical solutions.

Location: Palo Alto, CA, United States

Salary: $200,000–$400,000 USD annually, plus equity

Company

hirify.global is an infrastructure-first AI company building open systems for high-performance LLM inference and training, including SGLang and Miles.

What you will do

  • Profile and optimize GPU performance for production AI workloads across current and next-generation hardware.
  • Root-cause performance and correctness bottlenecks from GPU kernels through distributed multi-node systems.
  • Specialize in inference performance, GPU kernels and model enablement, speculative decoding, or training systems.
  • Build and optimize custom CUDA, ROCm, or Triton kernels and support new models on new hardware.
  • Partner with expert engineering teams to turn ambiguous technical problems into measurable improvements, cookbooks, and recommendations.
  • Feed user-driven improvements into the open-source SGLang and Miles systems and their future roadmap.

Requirements

  • 4+ years of experience in GPU systems, LLM infrastructure, or performance engineering.
  • Strong profiling, debugging, and performance root-cause analysis skills.
  • Hands-on experience with at least one of CUDA, ROCm, or Triton, with willingness to work across platforms.
  • Strong Python programming skills plus C++ or CUDA experience.
  • Ability to make progress on difficult, ambiguous problems and ramp quickly into unfamiliar systems and codebases.
  • Ability to translate ambiguous requests into clear technical plans, verified cookbooks, and actionable recommendations for expert engineering audiences.

Nice to have

  • Experience with LLM inference internals, including distributed serving, parallelism, routing, KV-cache management, and scheduling.
  • Experience with low-precision quantization such as FP8, INT8, or INT4, and custom GPU kernel optimization.
  • Familiarity with speculative decoding, large-scale distributed training, elasticity, or long-context workloads.
  • Experience optimizing across NVIDIA and AMD platforms.
  • Experience with SGLang, Miles, vLLM, TensorRT-LLM, Megatron, or comparable frameworks, including open-source AI/ML contributions.

Culture & Benefits

  • Work on open AI infrastructure intended to make frontier-level inference and training more accessible.
  • Collaborate with leading AI companies, research labs, and expert engineering partners.
  • Contribute to systems serving large-scale production workloads and distributed training environments.
  • Equity is included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →