Назад
Company hidden
5 часов назад

Member of Technical Staff — Inference-Kernel, Compiler & Communication (AI)

200 000 - 400 000$
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff — Inference-Kernel, Compiler & Communication (AI) (CUDA/C++/Python): Designing high-performance kernels, compiler and runtime optimizations, and distributed communication systems for frontier AI training and inference with an accent on GPU architecture, memory hierarchy, and large-scale accelerator clusters. Focus on reducing latency, increasing throughput, eliminating system bottlenecks, and building profiling tools for workloads spanning thousands of GPUs.

Location: Palo Alto, California, United States

Annual salary: $200,000–$400,000 USD plus equity

Company

hirify.global is an infrastructure-first AI company building open systems for large-scale inference and training, founded by AI infrastructure engineers from xAI and NVIDIA.

What you will do

  • Design and implement high-performance kernels for AI workloads.
  • Optimize compiler and runtime stacks for machine learning systems.
  • Improve communication efficiency across large GPU clusters.
  • Reduce latency, increase throughput, and eliminate bottlenecks across the systems stack.
  • Collaborate with training and inference teams on performance optimization.
  • Develop profiling tools and contribute to the architecture of performance-critical systems.

Requirements

  • 5+ years of experience in systems, compiler, or performance engineering.
  • Strong expertise in CUDA or accelerator programming and a deep understanding of GPU architecture and memory hierarchy.
  • Experience writing or optimizing high-performance kernels.
  • Strong background in compilers, runtimes, code generation, and distributed communication libraries such as NCCL, MPI, or RCCL.
  • Proficiency in C++ and Python, with strong system-level debugging and profiling skills.
  • Solid knowledge of networking and interconnect technologies.

Nice to have

  • Experience with Triton, TVM, XLA, MLIR, compiler passes, or IR transformations.
  • Familiarity with NVLink, InfiniBand, or RDMA.
  • Experience optimizing collective communication, scaling workloads to 1,000+ GPUs, or working with mixed-precision and quantized kernels.
  • Background in HPC or performance-critical systems and contributions to kernel, compiler, or ML systems open source.

Culture & Benefits

  • Work on low-level kernels, runtimes, compilers, and communication libraries for frontier AI systems.
  • Contribute to open infrastructure for inference and training.
  • Collaborate with engineers who have built production AI systems and large-scale GPU infrastructure.
  • Equity is included in the compensation package.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →