Match a job Paths Subjects Questions Quizzes Pricing

PPO & Modern Policy Optimization

Trust regions, the clipped surrogate objective, and Generalized Advantage Estimation — the algorithm behind RLHF

Overview Read

PPO & Modern Policy Optimization

Picture a warehouse-robotics team training a picking-arm policy with vanilla REINFORCE plus a learned baseline, the setup covered in the Policy Gradients & Actor-Critic subject. Training looks healthy for two hundred iterations — the success rate climbs from 40% to 81% — and then, on iteration 201, the success rate falls off a cliff to 6% and never recovers. Nothing in the code changed. No reward was misdefined. What happened is one unusually large, unusually confident gradient step: a minibatch where the advantage estimates happened to be large and consistent pushed the policy far enough that it started grasping air where a box used to be, generated a batch of terrible rollouts under the new policy, computed gradients from those terrible rollouts, and pushed itself further into a bad region it can no longer escape. The team cannot even roll back cleanly, because the optimizer's momentum state is now built on top of a collapsed policy.

This is the central failure mode that everything in this subject exists to prevent. In an interview, being able to explain why a single bad policy-gradient step is qualitatively more dangerous than a single bad supervised-learning step — and naming the two families of fixes, trust regions and clipped surrogates — is what separates a candidate who has read the PPO paper's abstract from one who has actually reasoned about on-policy RL. This subject builds that reasoning from the ground up: why unconstrained steps collapse policies, how TRPO constrains the step with a trust region, how PPO gets nearly the same guarantee with an objective you can optimize with plain SGD, how Generalized Advantage Estimation (GAE) supplies the advantage estimates PPO consumes, and the implementation details that decide whether any of this actually works in practice. It assumes you already have REINFORCE, baselines, the advantage function, and A2C from the Policy Gradients & Actor-Critic subject — this is the direct sequel.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.