Назад
2 дня назад

Midtraining Research Engineer (AI)

250 000 - 350 000$
Тип работы
fulltime
Грейд
middle
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Midtraining Research Engineer (AI) (LLM mid-training and scientific discovery): Improving frontier models' scientific reasoning by curating data, generating synthetic datasets, building evaluations, and running large-scale training experiments with an accent on self-distillation, on-policy distillation, and compute-efficient scaling. Focus on designing distributed training runs across thousands of GPUs, correlating evaluations with downstream scientific performance, and investigating how data choices shape model intelligence.

Location: Menlo Park, California; San Francisco planned soon

Compensation: $250,000–$350,000 per year plus equity

Company

Periodic Labs is an AI and physical sciences company developing frontier models to accelerate discoveries in materials, energy, and other scientific fields.

What you will do

  • Identify, process, and curate novel scientific data sources for large-scale model training.
  • Generate synthetic data to address gaps in scientific knowledge and reasoning.
  • Build evaluations that correlate with downstream scientific task performance.
  • Develop and apply self-distillation, on-policy distillation, and related model-improvement techniques.
  • Design and run large-scale training experiments across thousands of GPUs with supercompute engineers.
  • Build tools to investigate how data choices affect model intelligence.

Requirements

  • Experience training LLMs on curated mixes containing trillions of tokens.
  • Experience with mid-training or pre-training at scale.
  • Experience supporting large production training runs through dedicated evaluation work.
  • Hands-on experience with self-distillation, on-policy distillation, or similar methods in production training pipelines.
  • Ability to calculate scaling laws and compute-optimal hyperparameters.
  • Comfort working across data, evaluations, and training infrastructure; a bachelor's degree or equivalent experience is required.

Nice to have

  • Experience optimizing throughput and reliability for large-scale distributed training.
  • Background in AI for science or specialized scientific datasets such as protein or materials data.
  • Experience tracking evaluations and driving interventions during live large-scale training runs.
  • Experience with mid-training or pre-training at a major AI research lab.

Culture & Benefits

  • Work on frontier AI models for scientific discovery.
  • High-ownership environment focused on deep scientific expertise and rapid execution.
  • Equity included in the compensation package.
  • Visa sponsorship is available, with assistance throughout the process.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →