Advanced
Open
Pro
Justifying RLHF vs. DPO to a Skeptical Staff Engineer
Your team has collected 15,000 preference pairs (prompt, response A, response B, which one reviewers preferred) to make your assistant less prone to unnecessary hedging and more willing to give a direct answer when it has one. A staff engineer says: "RLHF is what all the big labs use for this kind of alignment work, so that's what we should build."
- Push back or agree — what would you actually recommend, and why?
- What does your recommended approach need infrastructurally that the other one wouldn't, and vice versa?
- Under what circumstances would the staff engineer's instinct (reach for RLHF) actually be right?
Share this question