Advanced
Open
Pro
Iterated Rejection-Sampling Fine-Tuning Drifting Toward the Easiest Solution Style
You run rejection-sampling fine-tuning in rounds: 1,000 questions, 64 samples each at nonzero temperature, keep every trace whose final answer verifies, fine-tune, repeat. In round 1, the 700 "easy" questions average 58 correct samples each — mostly near-identical traces — while the 300 "hard" questions average 1.2 correct samples each. After three rounds, aggregate pass rate is up, but pass rate on the hard slice is flat and the model's traces have become stylistically uniform.
- Quantify the round-1 training-set composition with no adjustments, and explain why iterating amplifies the imbalance rather than correcting it, even though each round's model is "better."
- The subject names two optional adjustments to RFT. Apply both to this pipeline, state what each fixes, and then identify a risk of the second adjustment that is specific to the hard slice — tie it to a weakness the subject attributes to outcome-only checking.
- Is this drift the same phenomenon as the STaR/RFT plateau the subject describes? Distinguish the two, and say which one moving to RL would actually address.
Share this question