Назад
13 дней назад

Staff AI Product Engineer (AI)

220 000 - 293 333$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Staff AI Product Engineer (AI): Setting the technical direction for inference serving, reinforcement learning, post-training systems, and developer APIs on a vertically integrated GenAI cloud platform with an accent on production LLM performance, GPU efficiency, and scalable platform architecture. Focus on designing serving and RL infrastructure, resolving latency, throughput, reliability, and cost challenges, and establishing reusable APIs and engineering standards across multiple teams.

Location: Houston, New York, San Francisco, or Seattle, United States

Salary: $220,000–$293,333 USD annually, plus potential bonus, equity, and/or commission.

Company

Nscale is building a vertically integrated GenAI cloud platform spanning data centers, software, and AI applications.

What you will do

  • Set the technical direction for LLM inference serving, including routing, scheduling, continuous batching, KV and prefix caching, speculative decoding, and model-efficiency strategies.
  • Own the architecture of reinforcement learning and post-training systems, including RLHF, DPO/GRPO-style methods, reward modelling, agentic RL, rollout generation, and policy updates.
  • Define how inference and training share GPU infrastructure and establish standards for fine-tuning, adapters, data curation, and processing workflows.
  • Design developer-facing APIs, SDKs, OpenAPI specifications, versioning, rate limiting, and reusable tooling for inference and RL capabilities.
  • Resolve systemic performance and reliability challenges across serving platforms, GPU fleets, multi-tenant environments, and distributed infrastructure.
  • Coach engineers, influence architecture across 2–4 teams, and align platform strategy with research, product, infrastructure, and customer needs.

Requirements

  • 8–12 years of engineering experience with significant depth in production AI systems at scale.
  • Deep expertise in production LLM inference, serving architectures, memory management, batching, scheduling, speculative decoding, and low-precision inference.
  • Strong hands-on experience with LLM reinforcement learning, preference optimization, reward modelling, or agentic and multi-turn RL on GPU clusters.
  • Experience designing developer APIs and SDKs, control-plane/data-plane architectures, and cell-based systems for scale and isolation.
  • Experience with GPU and accelerator workloads, CUDA or ROCm, memory optimization, distributed computing, parallelism, sharding, and scheduling.
  • Strong proficiency in Python and PyTorch, plus a deep understanding of transformer and LLM architectures under production load.

Nice to have

  • Contributions to inference or RL frameworks such as vLLM, SGLang, TensorRT-LLM, verl, OpenRLHF, TRL, or DeepSpeed.
  • Experience at an AI lab, hyperscaler AI team, or leading ML infrastructure company.
  • Published research or technical writing on inference systems or RL infrastructure.
  • Experience with custom CUDA kernels, Triton, Kubernetes, large-scale clusters, and model evaluation or benchmarking.

Culture & Benefits

  • Culture centered on innovation, ownership, accountability, openness, and transparency.
  • Collaboration across research, product, infrastructure, and engineering teams.
  • Medical, dental, and vision coverage.
  • Flexible paid time off, parental leave, and retirement plan participation.
  • Potential bonus, equity, and/or commission programs in addition to base salary.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →