Назад
3 дня назад

Systems Developer and Researcher (AI/LLM)

236 000 - 330 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Systems Developer and Researcher (AI/LLM): Developing high-performance LLM inference systems and optimization techniques with an accent on distributed serving, GPU kernel optimization, and model-system co-design. Focus on building intelligent, adaptive inference systems that automate performance tuning and push the boundaries of latency, throughput, and scalability.

Location: Bellevue, WA, USA

Compensation: $236,000 – $330,000 per year

Company

Snowflake is a cloud-based data platform company powering the era of the agentic enterprise through AI-native engineering.

What you will do

  • Design and develop high-performance LLM inference systems, including distributed serving and runtime systems.
  • Develop novel techniques to improve inference latency, throughput, memory efficiency, and cost.
  • Explore advanced inference techniques such as speculative decoding, KV-cache management, and adaptive parallelism.
  • Apply AI-native approaches to systems engineering, including automated profiling and bottleneck identification.
  • Develop strategies for multi-model serving, dynamic resource management, and model-system co-design.
  • Collaborate with research, infrastructure, and product teams to deploy innovations into production.

Requirements

  • Bachelor’s degree in Computer Science, Electrical Engineering, or related field (Master’s or PhD preferred).
  • 5+ years of experience in LLM inference systems, distributed AI systems, GPU systems, or HPC.
  • Strong understanding of modern LLM inference architectures and performance tradeoffs.
  • Hands-on experience with frameworks like vLLM, SGLang, or TensorRT-LLM.
  • Proficiency in GPU programming environments such as CUDA or Triton.
  • Experience profiling system performance using tools like Nsight Systems or Nsight Compute.

Nice to have

  • Experience using AI-native engineering approaches to accelerate software development and debugging.
  • Experience with performance libraries like CUTLASS, cuBLAS, or cuDNN.

Culture & Benefits

  • Opportunity to collaborate with world-class researchers and engineers from teams like DeepSpeed and vLLM.
  • Focus on experimental mindset and rapid testing of emerging capabilities.
  • Environment that encourages open-sourcing and publishing innovations in top-tier conferences.
  • Commitment to redefining the future of work through AI-native engineering.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →