Advanced
Open
Pro
Debug an Unstable PPO Run from Its Implementation Details
A junior engineer's PPO implementation for an inventory-restocking agent trains erratically: loss spikes every few hundred iterations, the policy sometimes reverts to near-random behavior mid-run, and performance is far below a published baseline on a similar task. Code review reveals: raw (unnormalized) advantages are fed directly into the clipped objective; the value function loss is a plain unclipped MSE; the rollout is reused for 20 full epochs of minibatch SGD before collecting new data; and there is no KL monitoring anywhere in the training loop.
- For each of the four issues, explain the specific mechanism by which it could cause the instability described.
- Which of the four would you fix first, and why?
- After fixing all four, what single scalar would you log every iteration to catch a future regression of this kind early, and what threshold behavior would you build around it?
Share this question