Advanced
Pro
Training Reasoning Models: STaR, RLVR, and Reward Models Quiz 2
A second pass through how reasoning models are trained: rejection-sampling fine-tuning vs. STaR, the plateau that motivates moving to RL, the full DeepSeek-R1 pipeline's stages, GRPO's literal advantage formula, how ORM training data is produced, why distillation skips RL entirely, what a Stream-of-Search training example actually contains, why RLVR concentrates in math and code, length-inflation mitigations, and what a passing test suite doesn't prove.
10 questions
15 min
Pass: 70%