Policy Gradients & Actor-Critic
Directly optimizing the policy — REINFORCE, baselines, advantage, and A2C
Learn the policy gradient theorem intuitively, REINFORCE as Monte Carlo policy gradient, why it has high variance, how a baseline and the advantage function reduce it without introducing bias, and how Advantage Actor-Critic (A2C) combines a policy (actor) with a learned value function (critic) into a stable, online training loop.
Practice questions (5)
-
View →
Compute a REINFORCE Update With and Without a Baseline
Advanced · Free -
View →
Choose an Advantage Estimator Under a Bias/Variance Constraint
Advanced -
View →
Debug a Collapsing Shared-Trunk Actor-Critic
Advanced -
View →
Evaluate a Proposal to Add a Replay Buffer to A2C
Advanced -
View →
Design a Policy Head for a Continuous Pricing Action
Advanced