Назад
Company hidden
2 дня назад

AI Researcher - Post-Training (f/m/d) (LLM)

Формат работы
onsite
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Germany
Релокация
Germany
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
AI Researcher - Post-Training (f/m/d) (LLM): Developing enterprise-grade coding agents and models through large language model post-training with an accent on policy optimization, verifier-based reinforcement learning, supervised fine-tuning, and synthetic data. Focus on designing research experiments, converting prototypes into products, and explaining advanced AI concepts to technical and non-technical audiences.

Location: Bochum, Germany; on-site. Candidates must be based in Bochum and attend the office, with anchor days on Mondays, Tuesdays, and Thursdays. Relocation support is available for the right candidate.

Company

hirify.global develops AI code verification, governance, and agentic software development products for enterprise customers.

What you will do

  • Develop and implement products that enable enterprise customers to post-train models for agentic coding practices.
  • Design research hypotheses and experiments with researchers, research engineers, MLOps specialists, and software engineers.
  • Iterate proofs of concept and convert successful prototypes into production-ready products.
  • Drive research into coding model post-training and contribute ideas as a subject matter expert.
  • Track advances in LLMs and agentic systems and explain complex technical concepts to technical and non-technical audiences.

Requirements

  • Master’s or PhD in Computer Science, Machine Learning, or a related quantitative field.
  • 4+ years of industry experience in machine learning and a strong understanding of modern software engineering practices.
  • Fluency in Python and experience with core machine learning frameworks.
  • Expertise in LLM post-training, including policy optimization algorithms such as GRPO and PPO, verifier-based reinforcement learning, supervised fine-tuning, synthetic data, data curation, parameter-efficient fine-tuning, and preference and safety alignment.
  • Experience leading research projects, delivering findings and prototypes, and converting them into products.
  • Excellent English communication skills and the ability to explain scientific topics clearly and concisely.

Nice to have

  • Experience with Rust or hirify.globalQube flagship languages such as C#, C++, JavaScript/TypeScript, or Java.

Culture & Benefits

  • Work in a cross-disciplinary team of researchers and engineers focused on practical, high-impact AI solutions.
  • Collaborate across global hubs in Austin, Bochum, Dubai, Geneva, London, Singapore, Tokyo, and Washington, D.C.
  • Relocation support is available for candidates who are willing to work in Bochum.
  • Employment is subject to a comprehensive background check and reference verification.
  • hirify.global is committed to diversity, equity, inclusion, and equal employment opportunity.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →