Назад
18 дней назад

ML Infrastructure Engineer (LLM)

200 000 - 287 500$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle/senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
ML Infrastructure Engineer (LLM): Building and scaling Cortex Training, a distributed platform for LLM post-training, inference, and reinforcement learning with an accent on multi-tenant GPU scheduling, orchestration, and fault-tolerant infrastructure. Focus on designing public APIs and SDKs, saturating GPU data planes under heavy concurrent load, and productionizing research techniques for enterprise-scale workloads.

Location: Bellevue, Washington, United States

Salary: $200,000–$287,500 per year

Company

Snowflake develops a cloud data platform, including ML infrastructure for running demanding AI workloads at enterprise scale.

What you will do

  • Design and build the full stack of Cortex Training, from public training APIs and SDKs to the control plane and GPU data plane.
  • Scale multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools.
  • Build fault-tolerant distributed systems for serverless GPU computing.
  • Optimize training, inference, and reinforcement-learning loops under heavy concurrent load while keeping GPUs saturated.
  • Partner with Snowflake Research to productionize state-of-the-art training and inference techniques.

Requirements

  • 3+ years of experience for the intermediate level or 6+ years for the senior level building and shipping production ML systems.
  • Strong distributed-systems and infrastructure foundation, including scalable, fault-tolerant services operated on Kubernetes in production.
  • Familiarity with GPU and LLM infrastructure such as PyTorch, DeepSpeed or FSDP, Ray, CUDA or NCCL, and vLLM.
  • Ability to debug across data, infrastructure, and GPU layers and harden systems for reliability, throughput, and cost efficiency.
  • Bachelor's degree in Computer Science or a related field.

Nice to have

  • Master's degree or PhD.
  • Hands-on LLM post-training or modeling experience.

Culture & Benefits

  • Work in an AI-focused engineering environment with an experimental approach to solving problems.
  • Collaborate with researchers and engineers on production ML infrastructure.
  • Build services for demanding enterprise-scale workloads.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →