STaR's Rationalization vs. Plain Rejection Sampling on a Hard Dataset
Your team is bootstrapping reasoning-trace training data for a model on a dataset of word problems where 70% are "easy" (the base model already gets these right most of the time when sampled) and 30% are "hard" (the base model almost never reaches the correct answer unassisted, even across many samples). Two options: (a) plain rejection-sampling fine-tuning (RFT) — sample many attempts per question, keep only ones with a verified-correct final answer; or (b) STaR with rationalization — same as (a), but for questions where no sampled attempt succeeds, show the model the correct answer and ask it to construct a plausible reasoning trace toward it.
- Predict what happens to the training set's composition under plain RFT (option a) specifically for the hard 30%, and why.
- Explain what rationalization changes about that outcome, and name the concrete risk it introduces that plain RFT does not have.
- Would you use rationalized traces at the same weight/frequency in training as genuinely-discovered correct traces? Justify your answer.
1. What happens under plain RFT on the hard 30%
Plain RFT can only keep an attempt if some sampled attempt actually reached the correct final answer. For the hard 30% where the base model "almost never" reaches the correct answer even across many samples, plain RFT will keep very few or zero training examples from that slice — the training set ends up dominated by the easy 70% (where correct attempts are common and cheap to collect), and the hard slice is effectively dropped or severely underrepresented. The resulting fine-tuned model improves on question types it could already mostly handle and gets little to no training signal on the question types it actually needed help with — the exact plateau this subject names as the reason to move beyond plain RFT/STaR toward RL, restated concretely for this dataset.
2. What rationalization changes, and its risk
Rationalization restores training coverage on the hard 30% by giving the model the correct answer and asking it to construct a plausible path to it, producing a training example even when the model could never find that path unassisted. The concrete risk this introduces that plain RFT does not have: a rationalized trace is constructed backward from a known answer, not discovered forward by actually reasoning through the problem — it can be a plausible-sounding but causally hollow justification (reasoning that looks locally coherent but doesn't reflect a real solving process, or quietly skips the actual hard step the model couldn't do unassisted and papers over it with confident-sounding connective text). Training on enough of these risks teaching the model to produce fluent-looking but non-substantive reasoning specifically on the hardest problem class — the opposite of what the intervention is meant to achieve.
3. Should rationalized traces be weighted the same as discovered ones?
No — the risk in part 2 is a direct reason to treat them differently, not identically. A defensible approach: keep rationalized traces as a smaller-weighted or explicitly flagged subset of the training data (so the training process, or a later audit, can distinguish "the model actually found this path" from "the model was shown the destination and asked to construct a road to it"), and specifically spot-check a sample of rationalized traces for the causally-hollow pattern described above before trusting them at scale. Weighting them identically risks quietly teaching the model that confident-sounding rationalization is an acceptable substitute for actually solving the hard 30% — precisely the failure this subject's "verify what a proxy is actually measuring" discipline (applied elsewhere to reward hacking) applies here to a training-data-quality question instead of a reward-signal one.
Share this question