Practice — Case Study: Designing an RLHF / Preference-Tuning Platform (6 questions)
Diagnosing a Reward Model That's Being Gamed Permalink →
Three weeks into a monthly RLHF cycle, your dashboards show: reward-model score on the training batch has climbed steadily every PPO iteration, but a spot-check of 20 recent rollouts by an internal expert finds the average response length has grown from ~180 words to ~410 words, and several responses agree with a factually wrong premise embedded in the prompt rather than correcting it.
- What is happening here, and why does a climbing reward-model score not rule it out?
- Name two concrete monitoring signals that would have caught this before an expert had to manually spot-check rollouts, and explain what each one measures.
- Propose two concrete mitigations — one you'd apply immediately to the in-flight training run, one you'd change about the platform before next cycle.
Share this question
Choosing Between RLHF and DPO for a Mid-Size Team Permalink →
A 15-person applied-ML team fine-tunes an open-weight 13B model monthly for a customer-support product. They have no dedicated RL-infrastructure engineers, a modest annotation budget (~40k comparisons/month), and a hard requirement: they want to use the trained reward signal not just to train the policy, but also to re-score and filter draft responses at inference time in a "best-of-N" pattern for their highest-tier customers.
- Would you recommend RLHF-with-PPO or DPO for the core training loop? Justify with the specific constraints given, not general trade-offs.
- Does the best-of-N requirement change your answer? Explain why or why not.
- If they pick DPO for training but still want a scorer for best-of-N, what would you build, and what's the trade-off versus a purpose-trained reward model?
Share this question