Назад
Company hidden
58 минут назад

Machine Learning Engineer, LLM Post-Training

150 000 - 230 000$
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Machine Learning Engineer, LLM Post-Training (LLM/RL): Leading continuous pre-training, supervised fine-tuning, and reinforcement learning for large language models, while building the data and evaluation pipelines that support product capabilities with an accent on RL methods, large-scale GPU training, and business-driven model development. Focus on designing preference and reward data, running distributed training with PyTorch and FSDP, and turning post-training research into production-ready systems.

Location: Mountain View, California, United States

Annual base salary: $150,000–$230,000 USD

Company

hirify.global is a content intelligence platform delivering personalized local news and information through AI, recommendation systems, and adtech.

What you will do

  • Lead LLM post-training across continuous pre-training, supervised fine-tuning, and reinforcement learning, with emphasis on RLHF, PPO, GRPO, DPO, and related methods.
  • Design and curate instruction, preference, reward, rollout, and rejection-sampled datasets for specific product and business scenarios.
  • Work with product and business stakeholders to translate use cases into training plans and targeted model capabilities.
  • Run large-scale training on mid-to-large GPU clusters using data parallelism, FSDP, and tensor or pipeline parallelism where appropriate.
  • Build evaluation, reward, and verifier pipelines to measure quality, prevent regressions, and maintain training–serving consistency.
  • Convert relevant post-training research into production-ready code.

Requirements

  • Hands-on experience personally running CPT, SFT, and RL training for LLMs, including practical experience with RLHF, PPO, GRPO, DPO, or similar methods.
  • Ability to independently design ML data-preparation strategies, including sourcing, cleaning, filtering, labeling, and synthetic or preference data generation.
  • Experience training LLMs on mid-to-large GPU hardware and debugging distributed training at scale.
  • Strong PyTorch fundamentals and working familiarity with Hugging Face TRL, Accelerate, DeepSpeed, FSDP, or vLLM.
  • Understanding of tokenization, attention, chat templates, and common alignment or agent-training failure modes.
  • Strong communication skills and a focus on rapid iteration and business impact.

Nice to have

  • Experience designing reward models or rule-based verifiers for RL.
  • Experience with tool-use or agentic model training, including function calling and multi-step planning.
  • Publications or open-source contributions in LLM post-training or RL.

Culture & Benefits

  • Health, dental, and vision coverage for employees and families, with 100% employee coverage.
  • 401(k) plan with company matching.
  • Paid time off and paid holidays.
  • FSA, HSA, and commuter benefits.
  • Team activity budget.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →