Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Choosing How to Build Self-Refinement Training Data When Correct Unassisted Traces Are Rare

You want to train a model to self-refine within one generation (catch and fix its own mistakes mid-trace, as a trained objective, not an inference-time harness loop). Two construction methods are on the table, both named in this subject: (a) mine the model's own inference-time generate-critique-revise transcripts, keeping only cases where revision measurably improved the outcome, and fine-tune on the full sequence as one training example; (b) use RL, with a reward that specifically credits a trajectory for catching and fixing its own earlier mistake. Your current model, unassisted, reaches a correct final answer only rarely on your target task distribution (similar to the "hard 30%" framing from STaR).

  1. Which construction method is more exposed to a data-scarcity problem given that "correct unassisted" is rare here, and why, tracing the mechanism precisely?
  2. The RL-reward version (b) has a reward-hacking-shaped risk specific to this training objective, distinct from the STaR rationalization risk. Name it precisely and explain the mechanism by which the reward could permit it.
  3. Propose one concrete way to detect the failure mode from part 2 in a trained model, distinct from just checking final-answer accuracy.

Share this question

← Back to Training Reasoning Models: STaR, RLVR, and Reward Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.