Practice — Training Reasoning Models: STaR, RLVR, and Reward Models (15 questions)
STaR's Rationalization vs. Plain Rejection Sampling on a Hard Dataset Permalink →
Your team is bootstrapping reasoning-trace training data for a model on a dataset of word problems where 70% are "easy" (the base model already gets these right most of the time when sampled) and 30% are "hard" (the base model almost never reaches the correct answer unassisted, even across many samples). Two options: (a) plain rejection-sampling fine-tuning (RFT) — sample many attempts per question, keep only ones with a verified-correct final answer; or (b) STaR with rationalization — same as (a), but for questions where no sampled attempt succeeds, show the model the correct answer and ask it to construct a plausible reasoning trace toward it.
- Predict what happens to the training set's composition under plain RFT (option a) specifically for the hard 30%, and why.
- Explain what rationalization changes about that outcome, and name the concrete risk it introduces that plain RFT does not have.
- Would you use rationalized traces at the same weight/frequency in training as genuinely-discovered correct traces? Justify your answer.
Share this question