Назад
Company hidden
5 дней назад

AI Infrastructure Engineer (GPU)

170 500 - 315 490$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Infrastructure Engineer (GPU): Optimizing LLM inference on Intel’s next-generation GPU architectures with an accent on custom kernels, cross-stack performance analysis, and open-source serving frameworks. Focus on designing high-performance attention, MoE, quantization, and operator-fusion kernels, upstreaming improvements to vLLM, SGLang, and PyTorch, and shaping future GPU roadmaps through real-world GenAI workload analysis.

Location: Hybrid work model in the United States, with on-site locations in Santa Clara and Folsom, California; Hillsboro, Oregon; or Austin, Texas.

Annual salary: $170,500–$315,490 USD.

Company

hirify.global develops semiconductor technologies, processors, GPUs, and platforms for computing and artificial hirify.globalligence.

What you will do

  • Own end-to-end performance optimization for state-of-the-art LLM inference on hirify.global GPUs.
  • Profile, diagnose, and resolve bottlenecks across the inference and systems stack.
  • Design and optimize custom GPU kernels for attention, mixture-of-experts, quantization, and operator fusion.
  • Upstream hardware backends and architectural improvements to vLLM, SGLang, PyTorch, and related open-source projects.
  • Apply roofline analysis and systematic profiling to guide GPU architecture and compiler roadmaps.
  • Collaborate with architecture, compiler, hardware, and open-source engineering teams.

Requirements

  • Bachelor’s degree in computer science, software engineering, artificial hirify.globalligence, machine learning, or a related field with 4+ years of experience; alternatively, a master’s degree with 3+ years or a PhD.
  • At least 3 years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing.
  • Strong proficiency in modern C++ and Python, including the ability to modify complex systems-level code.
  • Understanding of CPU/GPU architecture and modern LLM inference concepts, including attention, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation.
  • Hands-on experience with custom GPU kernels and technologies such as Triton, SYCL, CUDA, CUTLASS, or comparable domain-specific languages.
  • Ability to work in the United States under the stated hybrid work model.

Nice to have

  • Open-source contributions to vLLM, SGLang, PyTorch, llama.cpp, or other inference engines.
  • Experience with scale-out inference orchestration across multi-node topologies.
  • Experience using AI coding agents to accelerate development and benchmark generation.

Culture & Benefits

  • Hybrid schedule combining on-site work at an assigned hirify.global site with off-site work.
  • Competitive compensation with stock bonuses.
  • Health, retirement, and vacation benefits.
  • Focus on advancing AI infrastructure and hirify.global GPU technology.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →