Назад
Company hidden
4 дня назад

Senior Scientist - GenAI Evaluation

55 400 - 87 860€
Формат работы
hybrid
Тип работы
fulltime
Грейд
senior
Английский
b2
Страна
Spain
Вакансия из списка Hirify.GlobalВакансия из Hirify Global, списка международных tech-компаний
Для мэтча и отклика нужен Plus

Мэтч & Сопровод

Для мэтча с этой вакансией нужен Plus

Описание вакансии

Текст:
/
TL;DR
Senior Scientist - GenAI Evaluation (GenAI/Pharmaceutical R&D): Designing and operating evaluation methods for generative AI systems used in literature review, evidence synthesis, document Q&A, therapeutic-area search, and R&D decision support with an accent on scientific accuracy, traceability, safety, and reproducible quality measurement. Focus on building LLM and RAG evaluation pipelines, validating AI judges, curating benchmark datasets, and analyzing hallucinations and unsupported claims.

Location: Hybrid role based in Madrid or Barcelona, Spain; limited travel of less than 10%, primarily within Europe.

Base pay: €55,400–€87,860 per year.

Company

Johnson & Johnson MedTech develops healthcare and pharmaceutical R&D solutions, including generative AI systems for scientific workflows.

What you will do

  • Design, build, and maintain automated evaluation pipelines for LLM quality, RAG performance, agent reliability, safety, and scientific accuracy.
  • Author evaluation rubrics and scoring criteria, and curate golden and synthetic datasets with scientific domain experts.
  • Validate AI judges against human expert agreement and run model, prompt, retriever, and agent benchmarks.
  • Analyze hallucinations, unsupported claims, weak traceability, and other failure patterns to produce actionable recommendations.
  • Develop therapeutic-area-specific evaluation criteria with scientific, clinical, and regulatory partners.
  • Build reusable evaluation tooling that enables other teams to self-serve.

Requirements

  • Master's degree in AI/ML, Computer Science, Data Science, Computational Biology, Bioinformatics, Biomedical Engineering, Applied Mathematics, Biostatistics, or a related field; PhD preferred.
  • 6+ years of hands-on AI/ML evaluation or data science experience with a Master's degree, or 3+ years of industry experience with a PhD.
  • Experience designing evaluation frameworks, scientific benchmarks, or quality assessments for AI/ML systems.
  • Hands-on experience with generative AI, large language models, retrieval-augmented generation, agentic frameworks, and prompt engineering.
  • Strong proficiency in Python and modern AI/ML tooling, including evaluation harnesses, embedding models, vector databases, and LLM APIs.
  • English proficiency is required in written and verbal communication.

Nice to have

  • Experience evaluating AI/ML systems in regulated environments such as FDA- or EMA-related contexts.
  • Knowledge of drug development pipelines and biomedical data types, or expertise in oncology, immunology, or neuroscience.
  • Experience designing or validating LLM-as-judge systems, implementing CI/CD evaluation pipelines, or publishing in AI evaluation, NLP, or biomedical informatics.

Culture & Benefits

  • Work across multidisciplinary data science, scientific, clinical, and regulatory teams.
  • Annual bonus based on pay grade, location, and performance.
  • Vacation, parental leave of at least 12 weeks, bereavement leave, caregiver leave, and volunteer leave.
  • Well-being reimbursement and financial, physical, and mental health programs.
  • Insurance plans and service anniversary and recognition awards, subject to applicable plan terms.

Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →