Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

STaR's Rationalization vs. Plain Rejection Sampling on a Hard Dataset

Your team is bootstrapping reasoning-trace training data for a model on a dataset of word problems where 70% are "easy" (the base model already gets these right most of the time when sampled) and 30% are "hard" (the base model almost never reaches the correct answer unassisted, even across many samples). Two options: (a) plain rejection-sampling fine-tuning (RFT) — sample many attempts per question, keep only ones with a verified-correct final answer; or (b) STaR with rationalization — same as (a), but for questions where no sampled attempt succeeds, show the model the correct answer and ask it to construct a plausible reasoning trace toward it.

  1. Predict what happens to the training set's composition under plain RFT (option a) specifically for the hard 30%, and why.
  2. Explain what rationalization changes about that outcome, and name the concrete risk it introduces that plain RFT does not have.
  3. Would you use rationalized traces at the same weight/frequency in training as genuinely-discovered correct traces? Justify your answer.
Solution

1. What happens under plain RFT on the hard 30%

Plain RFT can only keep an attempt if some sampled attempt actually reached the correct final answer. For the hard 30% where the base model "almost never" reaches the correct answer even across many samples, plain RFT will keep very few or zero training examples from that slice — the training set ends up dominated by the easy 70% (where correct attempts are common and cheap to collect), and the hard slice is effectively dropped or severely underrepresented. The resulting fine-tuned model improves on question types it could already mostly handle and gets little to no training signal on the question types it actually needed help with — the exact plateau this subject names as the reason to move beyond plain RFT/STaR toward RL, restated concretely for this dataset.

2. What rationalization changes, and its risk

Rationalization restores training coverage on the hard 30% by giving the model the correct answer and asking it to construct a plausible path to it, producing a training example even when the model could never find that path unassisted. The concrete risk this introduces that plain RFT does not have: a rationalized trace is constructed backward from a known answer, not discovered forward by actually reasoning through the problem — it can be a plausible-sounding but causally hollow justification (reasoning that looks locally coherent but doesn't reflect a real solving process, or quietly skips the actual hard step the model couldn't do unassisted and papers over it with confident-sounding connective text). Training on enough of these risks teaching the model to produce fluent-looking but non-substantive reasoning specifically on the hardest problem class — the opposite of what the intervention is meant to achieve.

3. Should rationalized traces be weighted the same as discovered ones?

No — the risk in part 2 is a direct reason to treat them differently, not identically. A defensible approach: keep rationalized traces as a smaller-weighted or explicitly flagged subset of the training data (so the training process, or a later audit, can distinguish "the model actually found this path" from "the model was shown the destination and asked to construct a road to it"), and specifically spot-check a sample of rationalized traces for the causally-hollow pattern described above before trusting them at scale. Weighting them identically risks quietly teaching the model that confident-sounding rationalization is an acceptable substitute for actually solving the hard 30% — precisely the failure this subject's "verify what a proxy is actually measuring" discipline (applied elsewhere to reward hacking) applies here to a training-data-quality question instead of a reward-signal one.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.