Назад
5 дней назад

Midtraining Research Engineer (AI)

250 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Midtraining Research Engineer (AI): Improving frontier models' scientific reasoning by curating training data, generating synthetic data, building evaluations, and running large-scale training experiments with an accent on self-distillation, on-policy distillation, and scientific datasets. Focus on scaling training across thousands of GPUs, correlating evaluations with downstream scientific performance, and investigating how data choices shape model intelligence.

Location: Menlo Park, California, United States; on-site

Salary: $250,000–$350,000 per year plus equity

Company

Periodic Labs is an AI and physical sciences company developing models to accelerate breakthroughs in materials, energy, and scientific discovery.

What you will do

  • Identify, process, and curate scientific data for large-scale model training.
  • Generate synthetic data to address gaps in scientific knowledge and reasoning.
  • Build evaluations that correlate with downstream scientific task performance in collaboration with RL researchers, physicists, and chemists.
  • Develop self-distillation and on-policy distillation techniques to improve model capabilities.
  • Design and run large-scale training experiments across thousands of GPUs with supercompute engineers.
  • Build tools to investigate how data choices affect model intelligence.

Requirements

  • Experience training LLMs on curated mixtures containing trillions of tokens.
  • Experience supporting large production training runs through a dedicated evaluations team.
  • Hands-on experience with self-distillation, on-policy distillation, or similar methods in a real training pipeline.
  • Experience with scaling laws and compute-optimal hyperparameters.
  • Ability to work across data, evaluations, and training infrastructure.
  • Bachelor's degree or equivalent experience.

Nice to have

  • Experience optimizing throughput and reliability for large-scale distributed training.
  • Background in AI for science or training on specialized datasets such as protein or materials data.
  • Experience creating evaluations or synthetic data for non-verifiable tasks and tracking performance during live runs.

Culture & Benefits

  • Work alongside RL researchers, physicists, chemists, and supercompute engineers.
  • Operate in a rapidly growing environment focused on frontier scientific research.
  • Equity is included in the compensation package.
  • Visa sponsorship is available, with assistance throughout the process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →