Назад
6 дней назад

Research Engineer / Research Scientist, RL Frontiers (AI)

500 000 - 850 000$
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer / Research Scientist, RL Frontiers (AI): Developing and scaling reinforcement learning algorithms, model architectures, and experimental infrastructure for frontier-scale language-model training with an accent on throughput, stability, learning efficiency, and distributed systems. Focus on diagnosing scale-dependent failures, modeling compute and cost, and optimizing large RL runs from research code through accelerator hardware.

Location: San Francisco, New York City, or Seattle, United States; staff are expected to work from one of the offices at least 25% of the time.

Annual salary: $500,000–$850,000 USD

Company

Anthropic is a public benefit corporation developing reliable, interpretable, and steerable AI systems.

What you will do

  • Study how reinforcement learning training and sampling scale with model size, context length, and compute.
  • Develop next-generation transformer architectures and RL algorithms and run them efficiently at frontier scale.
  • Scale small-model results to large training runs and diagnose numerical, algorithmic, and systems-level differences.
  • Build experimental infrastructure for fast, reproducible comparisons of architecture and algorithm variants.
  • Own end-to-end performance of large RL runs, from research code to accelerator hardware.
  • Investigate instabilities, divergence, throughput regressions, and compute and cost trade-offs.

Requirements

  • Deep familiarity with transformer language models, their architectures, training dynamics, and large-scale optimization.
  • Hands-on experience training large models in distributed settings, including data, tensor, and pipeline parallelism.
  • Original technical work in ML training or systems demonstrated through research, open source, or production impact.
  • Ability to design rigorous large-scale experiments with baselines, ablations, and statistical validation.
  • Quantitative understanding of compute, memory, and communication costs.
  • Strong programming skills in Python and JAX or PyTorch, with the ability to modify code across the stack. A bachelor’s degree or equivalent experience is required.

Nice to have

  • Research experience in reinforcement learning, optimization, or large-scale training.
  • Experience developing RL algorithms for language models, scaling laws, or quantitative training-efficiency models.
  • Experience modifying transformer architectures and scaling training across large accelerator fleets.
  • Understanding of low-precision numerics, training instability, and GPU or TPU performance characteristics.
  • Experience with C++ or Rust.

Culture & Benefits

  • Collaborative research environment focused on large-scale efforts in safe and steerable AI.
  • Frequent research discussions and emphasis on communication and empirical science.
  • Flexible working hours and office-based collaboration.
  • Competitive compensation, optional equity donation matching, generous vacation, and parental leave.
  • Visa sponsorship may be available, with immigration-lawyer support, depending on the role and candidate.

Hiring process

  • Applications are evaluated against role-specific qualifications and internal job-level requirements.
  • Candidate AI usage is governed by Anthropic’s application policy.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →