Advanced
Open
Pro
Why Not Skip SFT and Go Straight From Base Model to DPO?
A team wants to save a training stage and go directly from a raw pretrained base model to DPO-style preference tuning, skipping SFT entirely, reasoning that "DPO already trains on preference pairs, so it should be able to teach the model to be helpful on its own."
- Explain concretely what would go wrong if you tried this, in terms of what the base model's outputs actually look like before SFT.
- What specific property does SFT give the model that preference tuning's comparative signal assumes already exists?
- Would this shortcut ever be defensible? Describe a narrow condition under which it might work reasonably, if any.
Share this question