Match a job Paths Subjects Questions Quizzes Pricing

Policy Gradients & Actor-Critic

Directly optimizing the policy — REINFORCE, baselines, advantage, and A2C

Overview Read

Policy Gradients & Actor-Critic

A programmatic ad platform bids on impressions in real time. For each auction, the agent chooses a continuous bid amount — not one of three or five discrete buckets, but any real number between $0.01 and the campaign's cap — based on the user, the publisher, the time of day, and the remaining daily budget. The Value-Based Methods: Q-Learning to DQN subject built agents that learn Q(s,a) and act by \arg\max_a Q(s,a); for a continuous bid, that \arg\max is itself an optimization problem to solve at every single auction, both during training and at serving time, with no closed form in general. Bidding is a small example of a much larger category: robot joint torques, portfolio allocation weights, dosage amounts, resource-allocation fractions — anywhere the natural action is a real number or a vector of real numbers, not a small discrete menu.

This subject covers the other branch of RL: policy-based methods, which skip the value function as an intermediary and directly parameterize and optimize the thing you actually want — the policy \pi_\theta(a \mid s) — by gradient ascent on expected return. Where value-based methods answer "how good is each action, so I can pick the best one," policy gradient methods answer "how should I nudge my action-choosing machinery to make good outcomes more likely." This distinction matters beyond elegance: a policy network can output the parameters of a continuous distribution (a mean and standard deviation for a Gaussian bid, say) just as naturally as it outputs a softmax over a handful of discrete actions, so policy gradients handle continuous and very large action spaces as a first-class case, exactly where value-based \arg\max search struggles or becomes intractable.

Value-based (Q-learning / DQN) Policy-based (REINFORCE / actor-critic)
What is learned Q(s,a), a value function \pi_\theta(a\mid s), the policy itself, directly
How action is chosen \arg\max_a Q(s,a) — derived from the value function Sampled directly from \pi_\theta
Continuous actions Requires solving an optimization at every step; awkward Natural — the policy outputs distribution parameters
Policy type Deterministic (from the argmax), unless post-hoc randomized Naturally stochastic — useful for exploration and for games with no deterministic optimum (e.g. rock-paper-scissors)
Sample reuse Off-policy variants (Q-learning) reuse old data via replay Classic REINFORCE/A2C are on-policy — each batch of data is used once, then discarded
Convergence behavior Can be unstable (deadly triad) but sample-efficient when stable More stable gradient-following, typically less sample-efficient

By the end of this subject you should be able to derive the REINFORCE gradient estimator, explain in one sentence why subtracting a baseline does not bias it, define the advantage function and say why it is the natural quantity a critic should estimate, and sketch the A2C training loop including the design choice between a shared and a separate actor/critic network.


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.