Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
ML Infrastructure Engineer (LLM): Building and scaling Cortex Training, a distributed platform for LLM post-training, inference, and reinforcement learning with an accent on multi-tenant GPU scheduling, orchestration, and fault-tolerant infrastructure. Focus on designing public APIs and SDKs, saturating GPU data planes under heavy concurrent load, and productionizing research techniques for enterprise-scale workloads.
Location: Bellevue, Washington, United States
Salary: $200,000–$287,500 per year
Company
Snowflake develops a cloud data platform, including ML infrastructure for running demanding AI workloads at enterprise scale.
What you will do
- Design and build the full stack of Cortex Training, from public training APIs and SDKs to the control plane and GPU data plane.
- Scale multi-tenant scheduling, placement, and capacity-aware routing across regional GPU pools.
- Build fault-tolerant distributed systems for serverless GPU computing.
- Optimize training, inference, and reinforcement-learning loops under heavy concurrent load while keeping GPUs saturated.
- Partner with Snowflake Research to productionize state-of-the-art training and inference techniques.
Requirements
- 3+ years of experience for the intermediate level or 6+ years for the senior level building and shipping production ML systems.
- Strong distributed-systems and infrastructure foundation, including scalable, fault-tolerant services operated on Kubernetes in production.
- Familiarity with GPU and LLM infrastructure such as PyTorch, DeepSpeed or FSDP, Ray, CUDA or NCCL, and vLLM.
- Ability to debug across data, infrastructure, and GPU layers and harden systems for reliability, throughput, and cost efficiency.
- Bachelor's degree in Computer Science or a related field.
Nice to have
- Master's degree or PhD.
- Hands-on LLM post-training or modeling experience.
Culture & Benefits
- Work in an AI-focused engineering environment with an experimental approach to solving problems.
- Collaborate with researchers and engineers on production ML infrastructure.
- Build services for demanding enterprise-scale workloads.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Runpod
3 дня назад
Senior ML Systems Engineer Inference (LLM)
150 000 - 220 000$
6 дней назад
Senior AI Infra Engineer (LLM/ML)
177 688 - 341 734$
4 дня назад
Research Platform Engineer (ML)
200 000 - 300 000$
5 дней назад
Senior ML Infrastructure & MLOps Engineer (Core Platform)
126 900 - 185 100$
4 дня назад
ML Platform Engineer (AI)
100 000 - 160 000$
5 дней назад
Senior AI/ML Platform Engineer
148 500 - 221 000$