Training Reasoning Models: STaR, RLVR, and Reward Models
How o-series-style and DeepSeek-R1-style models are actually trained — bootstrapped SFT on reasoning traces, RL against verifiable rewards, ORM vs. PRM, and the failure modes that come with all of it
The training-time half of the reasoning-model story for AI engineering interviews: SFT on model-generated reasoning traces (STaR, rejection-sampling fine-tuning); reinforcement learning with verifiable rewards (RLVR) and the DeepSeek-R1 recipe; GRPO vs. PPO at a glance — what GRPO drops and why; outcome vs. process reward models and auto-labeled process supervision without human step-level annotation; self-refinement as a training objective, distinct from inference-time self-refine; internalizing search (Meta-CoT, Stream of Search) so a model can search within one generation instead of needing external scaffolding; distilling reasoning traces into smaller models; named failure modes (overthinking, reward hacking, verifier coverage limits); and a worked reward-design example for a code-reasoning model using unit tests as the verifier.
Practice questions (15)
-
View →
STaR's Rationalization vs. Plain Rejection Sampling on a Hard Dataset
Advanced · Free -
View →
Would GRPO's Group-Relative Baseline Work for a Robotics RL Task?
Advanced -
View →
ORM or PRM for a Multi-Step Geometry Proof Assistant?
Advanced -
View →
Diagnosing Length Inflation vs. Reward Hacking in a Training Run
Advanced -
View →
Auditing a Reward Design for a Code-Reasoning Model
Advanced
Quizzes (3)
Training Reasoning Models: STaR, RLVR, and Reward Models Quiz 3
Training Reasoning Models: STaR, RLVR, and Reward Models Quiz 2
Training Reasoning Models: STaR, RLVR, and Reward Models Quiz