Назад
Company hidden
7 часов назад

Research Engineer – Interpretability Systems (AI)

150 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer – Interpretability Systems (AI): Building experimental systems for mechanistic interpretability and alignment research with an accent on activation tracing, internal representation probing, and model behavior analysis. Focus on detecting latent concepts, steering models at the activation level, and developing benchmarks for consistency and robustness.

Location: Onsite in San Francisco, CA

Salary: $150,000–$350,000 per year

Company

Early-stage, revenue-generating AI research lab working on interpretability, alignment, and reinforcement learning.

What you will do

  • Build experimental systems and custom tooling for mechanistic interpretability research.
  • Perform activation tracing and mechanistic analysis of large language models.
  • Probe internal model representations and detect latent concepts such as deception, goals, uncertainty, and hidden objectives.
  • Develop custom RL-style environments for alignment research.
  • Explore activation-level steering beyond prompting and fine-tuning.
  • Create benchmarks for model consistency and robustness.

Requirements

  • Strong software engineering fundamentals.
  • Experience with experimental ML or research systems.
  • Comfort working close to model internals.
  • Interest in interpretability, alignment, reinforcement learning, or mechanistic understanding.
  • Ability to work on ambiguous problems and operate through fast research cycles.
  • Onsite presence in San Francisco, CA is required.

Nice to have

  • PhD in a relevant field.

Culture & Benefits

  • Fast, experimental, greenfield research environment.
  • Build new tools from first principles and test research ideas rapidly.
  • Focus on research systems rather than production ML, MLOps, or large-scale training infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →