6 дней назад
Research Engineer / Research Scientist, RL Frontiers (AI)
500 000 - 850 000$
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
Текст:
TL;DR
Research Engineer / Research Scientist, RL Frontiers (AI): Developing and scaling reinforcement learning algorithms, model architectures, and experimental infrastructure for frontier-scale language-model training with an accent on throughput, stability, learning efficiency, and distributed systems. Focus on diagnosing scale-dependent failures, modeling compute and cost, and optimizing large RL runs from research code through accelerator hardware.
Location: San Francisco, New York City, or Seattle, United States; staff are expected to work from one of the offices at least 25% of the time.
Annual salary: $500,000–$850,000 USD
Company
Anthropic is a public benefit corporation developing reliable, interpretable, and steerable AI systems.
What you will do
- Study how reinforcement learning training and sampling scale with model size, context length, and compute.
- Develop next-generation transformer architectures and RL algorithms and run them efficiently at frontier scale.
- Scale small-model results to large training runs and diagnose numerical, algorithmic, and systems-level differences.
- Build experimental infrastructure for fast, reproducible comparisons of architecture and algorithm variants.
- Own end-to-end performance of large RL runs, from research code to accelerator hardware.
- Investigate instabilities, divergence, throughput regressions, and compute and cost trade-offs.
Requirements
- Deep familiarity with transformer language models, their architectures, training dynamics, and large-scale optimization.
- Hands-on experience training large models in distributed settings, including data, tensor, and pipeline parallelism.
- Original technical work in ML training or systems demonstrated through research, open source, or production impact.
- Ability to design rigorous large-scale experiments with baselines, ablations, and statistical validation.
- Quantitative understanding of compute, memory, and communication costs.
- Strong programming skills in Python and JAX or PyTorch, with the ability to modify code across the stack. A bachelor’s degree or equivalent experience is required.
Nice to have
- Research experience in reinforcement learning, optimization, or large-scale training.
- Experience developing RL algorithms for language models, scaling laws, or quantitative training-efficiency models.
- Experience modifying transformer architectures and scaling training across large accelerator fleets.
- Understanding of low-precision numerics, training instability, and GPU or TPU performance characteristics.
- Experience with C++ or Rust.
Culture & Benefits
- Collaborative research environment focused on large-scale efforts in safe and steerable AI.
- Frequent research discussions and emphasis on communication and empirical science.
- Flexible working hours and office-based collaboration.
- Competitive compensation, optional equity donation matching, generous vacation, and parental leave.
- Visa sponsorship may be available, with immigration-lawyer support, depending on the role and candidate.
Hiring process
- Applications are evaluated against role-specific qualifications and internal job-level requirements.
- Candidate AI usage is governed by Anthropic’s application policy.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →
Похожие вакансии
Thinking Machines Lab
9 дней назад
Mid-Training Researcher (AI)
350 000 - 475 000$
Figma
8 дней назад
AI Applied Scientist (AI)
153 000 - 376 000$
8 дней назад
Principal AI Researcher
228 000 - 342 000$
6 дней назад
Post-Training Engineer (AI)
300 000 - 350 000$
10 дней назад
Research Engineer (AI/ML)
Baseten
8 дней назад
Software Engineer (AI)
180 000 - 360 000$