1 день назад
Mid-Training Researcher (AI)
350 000 - 475 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Mid-Training Researcher (AI): Building and improving mid-training data pipelines, model recipes, quality systems, and evaluations for frontier AI models with an accent on synthetic data, knowledge injection, capability development, and scalable training. Focus on designing training datasets, calibrating quality classifiers and LLM judges, tuning large-model training recipes, and understanding how mid-training affects reasoning, coding, mathematics, and post-training stability.
Location: Hybrid role based in San Francisco, California, United States
Salary: $350,000–$475,000 annual salary
Company
Thinking Machines Lab builds AI systems intended to extend human judgment and communication, including frontier models, model-customization tools, and human-AI interfaces.
What you will do
- Own the sourcing, curation, synthesis, filtering, deduplication, verification, and rewriting of training data for model capabilities and knowledge areas.
- Design and measure knowledge-improvement interventions, including targeted corpora, synthetic rephrasings, question-answering over source documents, and knowledge-dense data mixes.
- Introduce and evaluate behaviors during mid-training and coordinate with post-training researchers on which behaviors belong in mid-training or reinforcement learning.
- Build automatic data-quality systems, including quality classifiers, LLM-based judges, synthetic-data verifiers, and fine-grained data-attribute controls.
- Develop and tune mid-training recipes covering dataset collections, training stages, annealing schedules, and hyperparameters.
- Define meaningful evaluations, debug unexpected training results, measure scaling behavior, and explore new training-data methodologies.
Requirements
- Proficiency in Python and familiarity with deep learning frameworks such as PyTorch, TensorFlow, or JAX.
- Experience debugging distributed training and writing code that scales.
- Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline.
- Strong theoretical and empirical grounding, with the ability to communicate complex technical concepts clearly in writing.
- Ability to work in a hybrid role based in San Francisco, California.
Nice to have
- Strong understanding of probability, statistics, and machine learning fundamentals.
- Experience building training datasets for large models, including synthetic-data pipelines, real-user data, collection, curation, filtering, or mixture design.
- Experience with knowledge injection, continued pre-training, domain adaptation, or measuring model knowledge.
- Experience building model-based data-quality systems, LLM judges, verification pipelines, or rewriting systems at billion-token scale.
- Research or engineering contributions in alignment, data-centric AI, or human-AI collaboration.
Culture & Benefits
- Individual contributor role within the Research Team, reporting to the Head of Research.
- Close collaboration with a small group of post-training researchers.
- Health, dental, and vision benefits.
- Unlimited paid time off and paid parental leave.
- Visa sponsorship and relocation support are available as needed.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
3 дня назад
AI Research Engineer
100 000 - 150 000$
5 дней назад
Machine Learning Leader (AI)
276 800 - 415 200$
2 дня назад
Director, AI & Machine Learning
242 543 - 328 147$
Resolution
6 дней назад
Research Engineer (AI Safety)
230 000 - 930 000$
6 дней назад
Research Engineer (AI/ML)
195 000 - 400 000$
2 дня назад