Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
LLM Post-training Engineer (AI): Designing and owning the RL and OPD post-training pipeline for a healthcare-focused LLM with an accent on clinical reasoning, safety, and alignment. Focus on building reward models, verifiers, and conversational AI simulations to ensure high-accuracy patient interactions.
Location: On-site in Palo Alto, CA (5 days a week)
Company
Hippocratic AI is a generative AI company developing the first safety-focused LLM specifically for healthcare to improve patient outcomes globally.
What you will do
- Design RL and OPD post-training methods such as RLHF and RLVR.
- Build and evaluate reward models, verifiers, and LLM-as-judge pipelines.
- Develop conversational AI environments and simulations for healthcare RL training using synthetic data.
- Automate post-training loops utilizing agents for auto-research.
- Execute rigorous experiments to determine drivers of post-training gains.
- Collaborate with research, engineering, and clinical teams.
Requirements
- MS or PhD in Computer Science or a relevant field.
- 3+ years of experience in NLP, LLM training, or RL.
- 1+ years of experience specifically in RL for LLM post-training.
- Experience with large-scale LLM training (50B+ parameters and multi-node).
- Strong proficiency in Python and PyTorch.
- Experience with RLHF, RLVR, LLM-as-judge, or similar post-training methods.
Culture & Benefits
- Work with a world-class team of AI pioneers, physicians, and researchers.
- Backed by leading investors including Avenir Growth, CapitalG, and a16z.
- Involved in category creation by building a world-first healthcare-only LLM.
- Fast-paced, collaborative environment with a strong on-site team culture.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →