Назад
Company hidden
6 дней назад

Member of Technical Staff - Compiler Engineer (AI)

180 000 - 400 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Member of Technical Staff - Compiler Engineer (AI) (MLIR/LLVM): Building and optimizing an AI compiler stack for executing machine-learning workloads across CPUs, GPUs, and emerging accelerators with an accent on intermediate representations, lowering paths, and heterogeneous hardware support. Focus on designing compiler passes, operator fusion, memory planning, code generation, and execution strategies that improve inference latency, throughput, utilization, and serving cost.

Location: San Francisco, CA; on-site

Salary: $180,000–$400,000 per year

Company

hirify.global is building a next-generation compute platform for fast and efficient AI inference across heterogeneous hardware.

What you will do

  • Build and scale an AI compiler stack based on MLIR and LLVM.
  • Design custom dialects, intermediate representations, optimization passes, and lowering paths for CPUs, GPUs, and emerging accelerators.
  • Develop high-performance code generation, operator fusion, tiling, layout transformation, memory planning, and execution-planning optimizations.
  • Optimize workload partitioning, device-to-device state movement, latency, throughput, utilization, and serving cost.
  • Integrate generated, library, and hand-written kernels into performance-driven code-generation strategies.
  • Collaborate with kernel, runtime, ML systems, distributed-systems, and infrastructure engineers.

Requirements

  • Experience building ML compilers or runtimes with a strong focus on performance optimization.
  • Experience with MLIR, LLVM, intermediate representations, compiler passes, lowering, and code-generation pipelines.
  • Strong C++ and/or Python skills.
  • Strong understanding of SSA, memory systems, scheduling, and hardware efficiency.
  • Experience with operator fusion, tiling, layout transformations, memory planning, and performance profiling.
  • Familiarity with IREE, XLA, TVM, Triton, or similar ML compiler and accelerator-programming frameworks.

Nice to have

  • Experience optimizing ML inference or model-serving workloads.
  • Kernel dispatch, runtime-interface development, memory allocators, or automated kernel generation.
  • Autotuning, search-based optimization, or contributions to open-source compiler infrastructure.
  • Experience with heterogeneous execution, distributed inference, speculative decoding, or prefill/decode disaggregation.

Culture & Benefits

  • Work on frontier models, production infrastructure, and emerging accelerator architectures.
  • Contribute to a platform supporting multiple heterogeneous compute targets.
  • Collaborate across compiler, kernel, runtime, ML systems, distributed-systems, and infrastructure disciplines.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →