Назад
Company hidden
4 дня назад

Research Engineer – Interpretability Systems (AI)

150 000 - 190 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer – Interpretability Systems (AI): Building experimental systems for mechanistic interpretability and alignment research with an accent on activation tracing, internal representation probing, and custom reinforcement-learning environments. Focus on detecting latent model concepts, steering models at the activation level, and developing benchmarks for consistency and robustness.

Location: San Francisco, CA — onsite

Salary: $150,000–$190,000 annually

Company

Early-stage, revenue-generating AI research lab working on interpretability, alignment, and reinforcement learning.

What you will do

  • Build experimental systems and custom tooling for interpretability research.
  • Perform activation tracing and mechanistic analysis of large language models.
  • Develop custom reinforcement-learning-style environments for alignment research.
  • Probe internal model representations and detect latent concepts such as deception, goals, uncertainty, and hidden objectives.
  • Develop activation-level steering methods beyond prompting and fine-tuning.
  • Create benchmarks for model consistency and robustness.

Requirements

  • Strong software engineering fundamentals.
  • Experience with experimental machine learning or research systems.
  • Comfort working close to model internals.
  • Interest in interpretability, alignment, reinforcement learning, or mechanistic understanding.
  • Ability to work on ambiguous problems and build new tools from first principles.

Nice to have

  • PhD in a relevant field.

Culture & Benefits

  • Fast, experimental, greenfield research environment.
  • Short research cycles focused on building tools, testing ideas, and evaluating results.
  • Focus on research systems rather than production ML, MLOps, or large-scale training infrastructure.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →