Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Why Not Skip SFT and Go Straight From Base Model to DPO?

A team wants to save a training stage and go directly from a raw pretrained base model to DPO-style preference tuning, skipping SFT entirely, reasoning that "DPO already trains on preference pairs, so it should be able to teach the model to be helpful on its own."

  1. Explain concretely what would go wrong if you tried this, in terms of what the base model's outputs actually look like before SFT.
  2. What specific property does SFT give the model that preference tuning's comparative signal assumes already exists?
  3. Would this shortcut ever be defensible? Describe a narrow condition under which it might work reasonably, if any.

Share this question

← Back to How LLMs Are Built: Pretraining to Chatbot practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.