Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Staff AI Product Engineer (AI): Setting the technical direction for inference serving, reinforcement learning, post-training systems, and developer APIs on a vertically integrated GenAI cloud platform with an accent on production LLM performance, GPU efficiency, and scalable platform architecture. Focus on designing serving and RL infrastructure, resolving latency, throughput, reliability, and cost challenges, and establishing reusable APIs and engineering standards across multiple teams.
Location: Houston, New York, San Francisco, or Seattle, United States
Salary: $220,000–$293,333 USD annually, plus potential bonus, equity, and/or commission.
Company
Nscale is building a vertically integrated GenAI cloud platform spanning data centers, software, and AI applications.
What you will do
- Set the technical direction for LLM inference serving, including routing, scheduling, continuous batching, KV and prefix caching, speculative decoding, and model-efficiency strategies.
- Own the architecture of reinforcement learning and post-training systems, including RLHF, DPO/GRPO-style methods, reward modelling, agentic RL, rollout generation, and policy updates.
- Define how inference and training share GPU infrastructure and establish standards for fine-tuning, adapters, data curation, and processing workflows.
- Design developer-facing APIs, SDKs, OpenAPI specifications, versioning, rate limiting, and reusable tooling for inference and RL capabilities.
- Resolve systemic performance and reliability challenges across serving platforms, GPU fleets, multi-tenant environments, and distributed infrastructure.
- Coach engineers, influence architecture across 2–4 teams, and align platform strategy with research, product, infrastructure, and customer needs.
Requirements
- 8–12 years of engineering experience with significant depth in production AI systems at scale.
- Deep expertise in production LLM inference, serving architectures, memory management, batching, scheduling, speculative decoding, and low-precision inference.
- Strong hands-on experience with LLM reinforcement learning, preference optimization, reward modelling, or agentic and multi-turn RL on GPU clusters.
- Experience designing developer APIs and SDKs, control-plane/data-plane architectures, and cell-based systems for scale and isolation.
- Experience with GPU and accelerator workloads, CUDA or ROCm, memory optimization, distributed computing, parallelism, sharding, and scheduling.
- Strong proficiency in Python and PyTorch, plus a deep understanding of transformer and LLM architectures under production load.
Nice to have
- Contributions to inference or RL frameworks such as vLLM, SGLang, TensorRT-LLM, verl, OpenRLHF, TRL, or DeepSpeed.
- Experience at an AI lab, hyperscaler AI team, or leading ML infrastructure company.
- Published research or technical writing on inference systems or RL infrastructure.
- Experience with custom CUDA kernels, Triton, Kubernetes, large-scale clusters, and model evaluation or benchmarking.
Culture & Benefits
- Culture centered on innovation, ownership, accountability, openness, and transparency.
- Collaboration across research, product, infrastructure, and engineering teams.
- Medical, dental, and vision coverage.
- Flexible paid time off, parental leave, and retirement plan participation.
- Potential bonus, equity, and/or commission programs in addition to base salary.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
14 дней назад
Staff Software Engineer, Product Experiences (AI)
200 000 - 275 000$
Anthropic
13 дней назад
Staff+ Software Engineer (Distributed Systems)
320 000 - 485 000$
14 дней назад
Staff AI Engineer
190 000 - 230 000$
13 дней назад
Staff AI Engineer (AI)
280 000 - 400 000$
14 дней назад
Software Engineering Intern (AI)
Scale AI
14 дней назад