Paths Subjects Questions Quizzes Pricing Search
Advanced Open Pro

Justifying RLHF vs. DPO to a Skeptical Staff Engineer

Your team has collected 15,000 preference pairs (prompt, response A, response B, which one reviewers preferred) to make your assistant less prone to unnecessary hedging and more willing to give a direct answer when it has one. A staff engineer says: "RLHF is what all the big labs use for this kind of alignment work, so that's what we should build."

  1. Push back or agree — what would you actually recommend, and why?
  2. What does your recommended approach need infrastructurally that the other one wouldn't, and vice versa?
  3. Under what circumstances would the staff engineer's instinct (reach for RLHF) actually be right?

Share this question

← Back to Fine-Tuning: SFT, LoRA/QLoRA, RLHF and DPO practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.