Назад
Company hidden
2 месяца назад

Research Kernel Engineer (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
Singapore/US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Kernel Engineer (CUDA/Triton): Developing and optimizing high-performance GPU kernels for novel AI quantization and sparsity methods with an accent on performance attribution and roofline analysis. Focus on turning research papers into production-viable kernels and maximizing hardware utilization for specific AI models.

Location: Singapore or Austin, US

Company

hirify.global is a world-leading technology company specializing in Bitcoin mining solutions and AI cloud infrastructure.

What you will do

  • Write and optimize CUDA and Triton kernels for new quantization schemes, sparsity patterns, and attention variants.
  • Conduct performance attribution across the serving path using roofline analysis to identify dominating kernels.
  • Convert published research methods into production-viable, high-performance implementations.
  • Develop the measurement basis and profiling frameworks used by the research team.
  • Collaborate closely with the AI Cloud platform team to integrate kernels into the serving stack.

Requirements

  • Degree in Computer Science, Electrical Engineering, or a related field.
  • Proficiency in CUDA, Triton, Python, and C++.
  • Deep understanding of GPU architecture and memory hierarchy, including occupancy and tensor cores.
  • Experience optimizing inference-critical paths such as attention, GEMM, and KV-cache management.
  • Ability to implement complex methods from research papers or whiteboard sketches.
  • Must be based in or able to work from Singapore or Austin, US

Nice to have

  • Experience with inference engines like vLLM, SGLang, or TensorRT-LLM.
  • Knowledge of compilers or IR-level tools such as MLIR, TVM, or TorchInductor.
  • Experience with distributed serving, including tensor and pipeline parallelism.
  • Publications at top-tier systems venues or significant open-source contributions to GPU computing.

Culture & Benefits

  • Inclusive environment with an exciting startup spirit and open workspaces.
  • Opportunity to network with industrial pioneers and contribute to the digital asset industry.
  • High degree of personal accountability, autonomy, and fast growth opportunities.
  • Attractive welfare benefits including training and mentoring.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →