Advanced
Open
Pro
Choosing a Search-Trace Source for Stream-of-Search Training Without Baking In Overthinking
You want to internalize search for a puzzle-like arithmetic planning task so no external tree-of-thought controller is needed at inference. The subject names two sources of training traces: (a) serialize an actual breadth-first search run over the state space, dead ends and all; (b) mine your own RLVR-trained model's traces that already show backtracking. Measured on your task: BFS traces average 4,000 tokens, the RL model's traces average 900, and a direct solution path averages 150.
- The subject presents overthinking as a failure of RL against an outcome-only reward. Explain mechanistically how SFT on source (a) can produce the same symptom with no RL involved, and why "just strip the dead ends before training" is not the fix.
- Compare (a) and (b) on what the model learns from each and on the plateau question: can source (b) contain search behavior for problems the RL model never solved? What does (a) give up in exchange for not having that ceiling?
- You plan an RLVR stage after the SoS SFT. Specify the reward shape and explain why adding an external tree-of-thought loop at inference afterward would likely be redundant.
Share this question