Advanced
Pro
Training Reasoning Models: STaR, RLVR, and Reward Models Quiz 3
A third, scenario-driven pass through reasoning-model training: computing a concrete GRPO advantage from a mixed group, the ORM blind spot for lucky-but-unsound traces, what Math-Shepherd rollout labels do and don't guarantee, telling reward hacking apart from a verifier coverage gap, scoring timeouts and crashes explicitly in a code reward, when an external self-refine loop becomes redundant, how distillation shifts the best-of-N trade-off, R1-Zero's emergent backtracking versus deliberately supervised Stream-of-Search, why RFT deduplicates near-identical traces, and what KL-anchoring actually buys against a proxy reward.
10 questions
15 min
Pass: 70%