Match a job Paths Subjects Questions Quizzes Pricing

Training Reasoning Models: STaR, RLVR, and Reward Models

How o-series-style and DeepSeek-R1-style models are actually trained — bootstrapped SFT on reasoning traces, RL against verifiable rewards, ORM vs. PRM, and the failure modes that come with all of it

Overview Read

Training Reasoning Models: STaR, RLVR, and Reward Models

reasoning-models-and-inference-time-scaling treats a reasoning model as a finished artifact — you're handed a model that adaptively allocates a thinking budget, and this subject's sibling covers how to spend more inference compute against it. That leaves the harder, more interview-differentiating question unanswered: how does a model become something that reasons well in the first place, rather than just being prompted to look like it does? This is where candidates who have only read model-card marketing copy stop, and where a strong AI engineering interview answer keeps going — into STaR's bootstrap loop, reinforcement learning against a reward you can actually verify, the specific thing GRPO drops relative to PPO and why that matters for language models specifically, and the failure modes (reward hacking, overthinking, a verifier that doesn't cover the case in front of it) that show up the moment any of this is deployed against a real task.

This subject assumes the RLHF/DPO mechanics from fine-tuning-sft-lora-rlhf-dpo — reward models trained on preference pairs, KL-anchored policy optimization — as background, and builds on top of them rather than re-deriving them; go there first if the reward-model-plus-policy-optimization shape is new. It also assumes the PPO mechanics (the clipped surrogate objective, GAE, the implementation details that make an RL run actually converge) from ppo-and-modern-policy-optimization, which this subject references by name rather than re-teaching — the goal here is to say precisely what GRPO changes relative to that baseline, not to re-derive the baseline itself. And it draws a hard, explicit line against its own sibling subject wherever the two could blur: reasoning-models-and-inference-time-scaling is about spending compute at answer time on a fixed model; everything below is about what changes in the model's weights to produce that model in the first place.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.