Мэтч & Сопровод
Для мэтча с этой вакансией нужен Plus
Описание вакансии
TL;DR
Senior AI Engineer (Healthcare): Own the Reinforcement Learning and On-Policy Distillation post-training pipeline for large language models to improve clinical reasoning, safety, and alignment in healthcare applications with an accent on RLHF, RLVR, and LLM-as-judge methods. Focus on designing post-training methods, building reward models, developing conversational AI environments, and running rigorous experiments to enhance model performance.
Location: On-site in Menlo Park, CA (Palo Alto office)
Company
Hippocratic AI is a leading generative AI startup focused on healthcare, building the world’s first safety-focused healthcare LLM with backing from top investors and experts.
What you will do
- Design RL and OPD post-training methods including RLHF, RLVR, and OPD.
- Build and evaluate reward models, verifiers, and LLM-as-judge pipelines.
- Develop conversational AI environments and simulations for healthcare RL training using synthetic data.
- Automate post-training loops with agents for auto-research.
- Run rigorous experiments to understand drivers of post-training gains.
- Collaborate closely with research, engineering, and clinical teams.
Requirements
- MS or PhD in Computer Science or relevant field.
- 5+ years experience in NLP, LLM training, or Reinforcement Learning.
- 2+ years experience specifically in RL for LLM post-training.
- Experience with large-scale (50B+ parameter, multi-node) LLM training.
- Strong Python and PyTorch coding skills.
- Experience with RLHF, RLVR, LLM-as-judge or similar post-training methods.
- Must be located on-site in Menlo Park, CA (Palo Alto office) five days a week.
Nice to have
- Publications at top AI/ML conferences (NeurIPS, ICML, ICLR, ACL, EMNLP).
- Healthcare domain experience.
Culture & Benefits
- Work alongside leading experts in healthcare and AI from top institutions and companies.
- Be part of a high-growth startup with significant funding and valuation.
- Collaborative, fast-paced environment emphasizing safety and innovation in healthcare AI.
Будьте осторожны: если работодатель просит войти в их систему, используя iCloud/Google, прислать код/пароль, запустить код/ПО, не делайте этого - это мошенники. Обязательно жмите "Пожаловаться" или пишите в поддержку. Подробнее в гайде →