Назад
Company hidden
3 дня назад

Senior ML Researcher (AI)

Формат работы
remote (только EU)/hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
UK/US/Serbia +2 еще
Вакансия из списка Hirify.GlobalВакансия из Hirify RU Global, списка компаний с восточно-европейскими корнями
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior ML Researcher (AI): Building end-to-end fine-tuning, reinforcement learning, and evaluation capabilities for an ML platform with an accent on LLM post-training, reward modeling, and scalable production workflows. Focus on designing GRPO-style methods, calibrating LLM-as-judge systems against human labels, and turning applied research experiments into self-serve platform features.

Location: European Union. Full remote is available; hybrid work is available if based in the Netherlands or Serbia.

Company

hirify.global AI builds data and machine-learning products that help leading AI companies train, evaluate, and improve generative AI models.

What you will do

  • Own end-to-end fine-tuning pipelines, including data preparation, SFT/LoRA training, model distillation, evaluation, and serving handoff.
  • Extend the post-training stack with reinforcement learning, GRPO-style methods, reward modeling, and LLM-judge-based rewards.
  • Build evaluation harnesses calibrated against human labels, including golden datasets and regression evaluations.
  • Improve the platform’s guiding agent through prompt and tool design, evaluation-driven iteration, and stress testing.
  • Run experiments for client projects and convert successful approaches into repeatable platform capabilities.
  • Collaborate with platform engineers on SDK and API interfaces for training and evaluation workflows.

Requirements

  • 4+ years of experience in ML engineering or applied research, including 1–2 years working hands-on with LLMs.
  • Practical experience fine-tuning open-weight models with LoRA or full fine-tuning, including data curation.
  • Strong understanding of LLM evaluation and LLM-as-judge calibration against human judgments.
  • Strong Python engineering skills and experience with PyTorch, the Hugging Face ecosystem, and vLLM or similar tools.
  • Product mindset and ability to work effectively with ambiguity and changing priorities.
  • English fluency at B2 level or above; candidates must be based in the European Union.

Nice to have

  • Experience with RLHF, RLAIF, GRPO, PPO-style post-training, or reward modeling.
  • Experience with DSPy, GEPA, prompt compression, distillation, or quantization.
  • Experience building agentic systems or shipping ML features in self-serve products.

Culture & Benefits

  • International, remote-first environment with a globally distributed team.
  • Full remote or hybrid work model, with hybrid availability in the Netherlands and Serbia.
  • Competitive compensation package with base salary, bonus, and ESOP.
  • Paid PTO, location-dependent benefits, IT setup, and home office allowances.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →