Назад
обновлено 8 дней назад

Research Mid-Training (AI)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Mid-Training (AI): Own late-stage training decisions sharpening raw base model capabilities into reliable reasoning foundations with an accent on data mix, quality uplift, annealing schedules, context length extension, and synthetic data strategies. Focus on capability injection across coding, math, and long-horizon reasoning, evaluating interventions, and scaling methodologies for AI agents like Devin.

Location: San Francisco, United States; on-site

Company

Building Devin, an AI software engineer, and developing models that can reason about real-world tasks.

What you will do

  • Design and iterate on high-quality data mixtures for late-stage and annealing training runs.
  • Drive capability improvements in coding, mathematics, and long-horizon reasoning through data strategies and training interventions.
  • Develop and evaluate synthetic data pipelines that generate training signals at scale.
  • Research learning-rate schedules, warmup strategies, compute allocation, and context-length extension methods.
  • Build evaluations that distinguish genuine capability improvements from benchmark overfitting.
  • Measure how mid-training interventions scale with compute and data, developing new methods when existing approaches reach their limits.

Requirements

  • Deep familiarity with the end-to-end LLM training pipeline, including pre-training data, optimization, architecture, and interactions between mid-training and post-training.
  • Hands-on experience with continual pre-training, annealing, or late-stage data mixing for large models.
  • Experience developing or evaluating synthetic data pipelines for capability improvement.
  • Proficiency in Python and PyTorch, with the ability to debug distributed training at scale.
  • Strong fundamentals in optimization, statistics, and machine learning theory, plus a track record of original contributions.
  • Ability to work in ambiguous, fast-moving environments and prioritize demonstrated capability over credentials.

Culture & Benefits

  • Small, highly selective team where research and product development move together.
  • Research prototypes reach real deployment quickly.
  • Large compute allocations with training jobs routinely running across thousands of GPUs.
  • Environment emphasizing speed, autonomy, technical depth, and minimal process overhead.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →