Назад
Company hidden
обновлено 1 день назад

Research Lead / Principal Scientist & Manager Post-Training (AI)

Формат работы
remote (Global)
Тип работы
fulltime
Грейд
senior/lead
Английский
c1
Страна
UK/US/Canada
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Research Lead / Principal Scientist & Manager Post-Training (AI): Designing and leading the research strategy for transforming foundation models into reliable, domain-specific systems with an accent on RLHF, preference optimization, and long-horizon reasoning. Focus on grounding reinforcement learning in physical laws and CAD kernels to ensure robustness in high-precision engineering domains.

Location: Remote (US, Canada, EU) or Toronto, ON, Canada

Company

hirify.global AI Lab advances state-of-the-art research across generative AI, multimodal foundation models, and reasoning systems to impact industries that shape the physical world.

What you will do

  • Own post-training strategy for model development, covering RLHF, preference optimization, agentic systems, and long-horizon reasoning.
  • Develop novel algorithms to improve model reliability, controllability, and alignment.
  • Manage and mentor a growing team of AI scientists, setting technical direction and research priorities.
  • Design evaluation frameworks for tool use, agentic behavior, and real-world workflow completion.
  • Partner with infrastructure teams to build scalable and reproducible post-training workflows.
  • Contribute to publications at top-tier venues (NeurIPS, ICML, ICLR, CVPR, SIGGRAPH) and patents.

Requirements

  • Deep hands-on expertise in reinforcement learning for foundation models (RLHF, RLAIF, DPO, PPO).
  • Proven experience leading or mentoring technical research teams in academia or industry.
  • PhD or equivalent depth of industry research experience in ML, RL, AI, or a related field.
  • Strong intuition for model behavior, alignment challenges, and post-training trade-offs.
  • Ability to communicate complex technical trade-offs to both technical and non-technical audiences.
  • Must be based in the US, Canada, or the EU.

Nice to have

  • Experience at a frontier model lab or advanced applied AI organization.
  • Strong publication record at leading ML or AI venues.
  • Background in alignment research, preference learning, or agentic AI.
  • Experience deploying or supporting production AI systems.

Culture & Benefits

  • Collaborative environment with a direct line from research advances to product impact at scale.
  • Access to unique, domain-grounded verifiers based on physics simulation and CAD kernels.
  • Inclusive culture committed to diversity and equal opportunity.
  • Competitive compensation package including annual cash bonuses and stock grants.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →