Research Scientist (AI Safety)
Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Research Scientist (AI Safety): Developing post-training methods and interpretability techniques to make frontier AI systems safer and more transparent with an accent on model robustness and alignment properties. Focus on designing post-training pipelines, creating interpretability-informed evaluations, and translating research into actionable safety standards.
Location: San Francisco, CA; New York, NY; Seattle
Salary: $216,000 - $270,000 USD
Company
is a leading data and evaluation partner for frontier AI companies, focusing on providing high-quality data and technology to develop reliable AI systems for critical decisions.
What you will do
- Design and run post-training pipelines to study how training choices affect model safety, robustness, and alignment.
- Develop interpretability-informed evaluations to reveal and mitigate unsafe, deceptive, or undesirable model behaviors.
- Collaborate with policymakers and engineers to translate research findings into safety standards and benchmarks.
- Tackle complex problems in agent robustness, AI control protocols, and AI risk evaluations.
- Publish findings to help governments and industry understand and mitigate AI risks.
Requirements
- Experience with post-training and RL techniques such as RLHF, DPO, GRPO.
- Track record of published research in machine learning, specifically in generative AI.
- At least three years of experience addressing sophisticated ML problems in research or product development.
- Strong written and verbal communication skills for operating in cross-functional teams.
- Deep commitment to promoting safe, secure, and trustworthy AI deployments.
Nice to have
- Experience with mechanistic interpretability, probing, or other model internal analysis techniques.
- Familiarity with red-teaming or adversarial evaluation of post-trained models.
- Experience studying failure modes like reward hacking, sycophancy, or alignment faking.
Culture & Benefits
- Comprehensive health, dental, and vision coverage.
- Equity compensation based on Board of Director approval.
- Retirement benefits and a learning and development stipend.
- Generous PTO and commuter stipend for eligible roles.
Hiring process
- Research interviews focused on practical ML prototyping and debugging.
- Assessment of research concepts and cultural alignment.
- No LeetCode-style questions used in the evaluation.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →