Advanced
Pro
Training Reasoning Models: STaR, RLVR, and Reward Models Quiz
Test your understanding of how reasoning models are actually trained: STaR and rejection-sampling fine-tuning, RL with verifiable rewards (RLVR) and the DeepSeek-R1 recipe, GRPO vs. PPO, ORM vs. PRM and auto-labeled process supervision, training-time self-refinement, internalizing search, distillation, and the named failure modes (overthinking, reward hacking, verifier coverage limits).
11 questions
15 min
Pass: 70%