5 дней назад
AI Infrastructure Engineer (GPU)
170 500 - 315 490$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
AI Infrastructure Engineer (GPU): Optimizing LLM inference on Intel’s next-generation GPU architectures with an accent on custom kernels, cross-stack performance analysis, and open-source serving frameworks. Focus on designing high-performance attention, MoE, quantization, and operator-fusion kernels, upstreaming improvements to vLLM, SGLang, and PyTorch, and shaping future GPU roadmaps through real-world GenAI workload analysis.
Location: Hybrid work model in the United States, with on-site locations in Santa Clara and Folsom, California; Hillsboro, Oregon; or Austin, Texas.
Annual salary: $170,500–$315,490 USD.
Company
develops semiconductor technologies, processors, GPUs, and platforms for computing and artificial ligence.
What you will do
- Own end-to-end performance optimization for state-of-the-art LLM inference on GPUs.
- Profile, diagnose, and resolve bottlenecks across the inference and systems stack.
- Design and optimize custom GPU kernels for attention, mixture-of-experts, quantization, and operator fusion.
- Upstream hardware backends and architectural improvements to vLLM, SGLang, PyTorch, and related open-source projects.
- Apply roofline analysis and systematic profiling to guide GPU architecture and compiler roadmaps.
- Collaborate with architecture, compiler, hardware, and open-source engineering teams.
Requirements
- Bachelor’s degree in computer science, software engineering, artificial ligence, machine learning, or a related field with 4+ years of experience; alternatively, a master’s degree with 3+ years or a PhD.
- At least 3 years of relevant software engineering experience in GPU computing, AI systems, or high-performance computing.
- Strong proficiency in modern C++ and Python, including the ability to modify complex systems-level code.
- Understanding of CPU/GPU architecture and modern LLM inference concepts, including attention, KV caching, continuous batching, speculative decoding, and prefill-decode disaggregation.
- Hands-on experience with custom GPU kernels and technologies such as Triton, SYCL, CUDA, CUTLASS, or comparable domain-specific languages.
- Ability to work in the United States under the stated hybrid work model.
Nice to have
- Open-source contributions to vLLM, SGLang, PyTorch, llama.cpp, or other inference engines.
- Experience with scale-out inference orchestration across multi-node topologies.
- Experience using AI coding agents to accelerate development and benchmark generation.
Culture & Benefits
- Hybrid schedule combining on-site work at an assigned site with off-site work.
- Competitive compensation with stock bonuses.
- Health, retirement, and vacation benefits.
- Focus on advancing AI infrastructure and GPU technology.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
5 часов назад
Inference Engineer (AI)
195 000 - 285 000$
2 часа назад
AI Infrastructure Engineer (LLM)
250 000 - 300 000$
1 минуту назад
Senior Performance Engineer (AI Inference)
5 часов назад
Systems/GPU Engineer (AI)
160 000 - 320 000$
4 часа назад
Senior Software Engineer, AI Infrastructure
126 000 - 189 000$
4 часа назад
Principal System Software Engineer, AI Inference Execution
195 000 - 285 000$