Match a job Paths Subjects Questions Quizzes Pricing
Machine Learning Advanced Pro

Policy Gradients & Actor-Critic

Directly optimizing the policy — REINFORCE, baselines, advantage, and A2C

30 min read 19 views

Learn the policy gradient theorem intuitively, REINFORCE as Monte Carlo policy gradient, why it has high variance, how a baseline and the advantage function reduce it without introducing bias, and how Advantage Actor-Critic (A2C) combines a policy (actor) with a learned value function (critic) into a stable, online training loop.

Practice questions (5)

  • Compute a REINFORCE Update With and Without a Baseline

    Advanced · Free
    View →
  • Choose an Advantage Estimator Under a Bias/Variance Constraint

    Advanced
    View →
  • Debug a Collapsing Shared-Trunk Actor-Critic

    Advanced
    View →
  • Evaluate a Proposal to Add a Replay Buffer to A2C

    Advanced
    View →
  • Design a Policy Head for a Continuous Pricing Action

    Advanced
    View →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.