Назад
Company hidden
5 часов назад

Research Engineer, Interpretability Systems

250 000 - 350 000$
Формат работы
onsite
Тип работы
fulltime
Английский
b2
Страна
US
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Engineer, Interpretability Systems (AI/Deep Learning): Building experimental infrastructure, RL-style environments, and tooling for mechanistic interpretability and alignment research in large language models with an accent on model internals, activation tracing, concept detection, and activation-level steering. Focus on implementing probes for latent concepts, defining robustness benchmarks, and rapidly turning research ideas into experiments and measurable results.

Location: San Francisco Bay Area, on-site

Salary: $250,000–$350,000 per year

Company

Early-stage AI research lab founded by former frontier-model researchers and focused on alignment and interpretability for large language models.

What you will do

  • Build custom RL-style environments and experimental testbeds for interpretability research.
  • Develop tooling for activation tracing, concept detection, and mechanistic analysis of model representations.
  • Implement probes for latent concepts such as deception, uncertainty, goals, and hidden objectives.
  • Prototype activation-level steering methods beyond prompting and fine-tuning.
  • Collaborate with researchers to iterate from research ideas to implementations, experiments, and results.
  • Define benchmarks and measurement frameworks for model consistency, robustness, alignment, and interpretability.

Requirements

  • Strong software engineering fundamentals and experience building experimental machine learning systems.
  • Experience working with model internals, representations, or post-training systems.
  • Strong Python skills and experience with deep learning frameworks, preferably PyTorch.
  • Ability to rapidly prototype and iterate in open-ended research environments.
  • Experience in interpretability, alignment, or machine learning research is preferred.
  • A PhD is considered ideal.

Culture & Benefits

  • Early-stage research environment focused on cutting-edge alignment and interpretability work.
  • Close collaboration between research engineers and researchers.
  • Opportunity to build infrastructure enabling new classes of model understanding and control experiments.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →