Назад
Company hidden
6 часов назад

Research Engineer - AI Performance & Kernel Optimization

Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer - AI Performance & Kernel Optimization (AI systems, GPU kernels): Improving large-scale language model training and inference stacks through optimized kernels, accelerator performance tuning, and distributed systems engineering with an accent on CUDA, HIP, Triton, and non-NVIDIA hardware. Focus on profiling memory, communication, scheduling, and compute bottlenecks, optimizing large MoE model parallelism, and translating low-level improvements into gains in throughput, latency, and hardware utilization.

Location: On-site in San Francisco, United States

Company

hirify.global develops frontier-scale AI systems and large-scale language model training and inference infrastructure.

What you will do

  • Develop and optimize kernels for large-scale machine learning workloads using PTX, assembly, CUDA, HIP, Triton, or other GPU DSLs.
  • Tune training and inference stacks across GPUs and other accelerator platforms.
  • Profile and eliminate bottlenecks in memory movement, communication, scheduling, compute utilization, and kernel execution.
  • Optimize distributed training and inference for large mixture-of-experts models, including model and parallelism strategies.
  • Improve portability and performance on non-NVIDIA hardware, including AMD MI300x and MI355x accelerators.
  • Collaborate with research and infrastructure teams to translate systems improvements into practical model training and inference gains.

Requirements

  • Experience writing highly performant GPU kernels with PTX, CUDA, HIP, Triton, or another kernel DSL.
  • Experience optimizing machine learning workloads for large-scale training, language model pretraining, or inference.
  • Strong understanding of distributed training and parallelism, including data, tensor/model, pipeline parallelism, sharding, and communication/computation overlap.
  • Strong systems intuition covering memory hierarchy, bandwidth constraints, kernel fusion, launch overhead, communication overhead, and hardware utilization.
  • Experience with profiling and debugging tools, collective communication libraries, and runtime performance analysis.
  • Background in physics, mathematics, theoretical computer science, computer science, electrical engineering, or another highly technical field.

Nice to have

  • Experience with AMD, AWS Trainium, Google TPU, Qualcomm, ARM, Intel, or custom ASIC hardware.
  • Performance engineering experience in HPC, quantitative finance, scientific computing, graphics, compilers, or numerical simulation.
  • HPC experience.

Culture & Benefits

  • Research and engineering excellence are valued through methodical, step-by-step work toward ambitious goals.
  • New and unconventional ideas are encouraged, with a willingness to invest in high-potential approaches.
  • In-person collaboration in a high-energy San Francisco environment.
  • Medical, dental, vision, and FSA plans, plus a 401(k) plan.
  • Unlimited PTO, company holidays, in-office snacks and meals, and case-by-case relocation and immigration support.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →