Choosing How to Build Self-Refinement Training Data When Correct Unassisted Traces Are Rare
You want to train a model to self-refine within one generation (catch and fix its own mistakes mid-trace, as a trained objective, not an inference-time harness loop). Two construction methods are on the table, both named in this subject: (a) mine the model's own inference-time generate-critique-revise transcripts, keeping only cases where revision measurably improved the outcome, and fine-tune on the full sequence as one training example; (b) use RL, with a reward that specifically credits a trajectory for catching and fixing its own earlier mistake. Your current model, unassisted, reaches a correct final answer only rarely on your target task distribution (similar to the "hard 30%" framing from STaR).
- Which construction method is more exposed to a data-scarcity problem given that "correct unassisted" is rare here, and why, tracing the mechanism precisely?
- The RL-reward version (b) has a reward-hacking-shaped risk specific to this training objective, distinct from the STaR rationalization risk. Name it precisely and explain the mechanism by which the reward could permit it.
- Propose one concrete way to detect the failure mode from part 2 in a trained model, distinct from just checking final-answer accuracy.
Share this question