Advanced
Open
Pro
Map PPO's Components onto an RLHF Fine-Tuning Pipeline
An interviewer asks: "We use PPO to fine-tune our chat model with RLHF. Walk me through what changes and what stays the same compared to using PPO for a game-playing agent, and where a naive application of vanilla policy gradient would go wrong here specifically."
- Map "policy," "reward," and "trajectory/episode" onto the RLHF setting, and explain how each differs structurally from a multi-step game agent.
- Explain what the KL-to-reference term does in RLHF PPO, and how it relates to (but differs from) TRPO's KL constraint.
- Describe, concretely, what "policy collapse" looks like in the RLHF setting, and why it is arguably more costly to let happen than in a simulated game environment.
Share this question