Policy Gradients & Actor-Critic
A programmatic ad platform bids on impressions in real time. For each auction, the agent chooses a continuous bid amount — not one of three or five discrete buckets, but any real number between $0.01 and the campaign's cap — based on the user, the publisher, the time of day, and the remaining daily budget. The Value-Based Methods: Q-Learning to DQN subject built agents that learn Q(s,a) and act by \arg\max_a Q(s,a); for a continuous bid, that \arg\max is itself an optimization problem to solve at every single auction, both during training and at serving time, with no closed form in general. Bidding is a small example of a much larger category: robot joint torques, portfolio allocation weights, dosage amounts, resource-allocation fractions — anywhere the natural action is a real number or a vector of real numbers, not a small discrete menu.
This subject covers the other branch of RL: policy-based methods, which skip the value function as an intermediary and directly parameterize and optimize the thing you actually want — the policy \pi_\theta(a \mid s) — by gradient ascent on expected return. Where value-based methods answer "how good is each action, so I can pick the best one," policy gradient methods answer "how should I nudge my action-choosing machinery to make good outcomes more likely." This distinction matters beyond elegance: a policy network can output the parameters of a continuous distribution (a mean and standard deviation for a Gaussian bid, say) just as naturally as it outputs a softmax over a handful of discrete actions, so policy gradients handle continuous and very large action spaces as a first-class case, exactly where value-based \arg\max search struggles or becomes intractable.
| Value-based (Q-learning / DQN) | Policy-based (REINFORCE / actor-critic) | |
|---|---|---|
| What is learned | Q(s,a), a value function | \pi_\theta(a\mid s), the policy itself, directly |
| How action is chosen | \arg\max_a Q(s,a) — derived from the value function | Sampled directly from \pi_\theta |
| Continuous actions | Requires solving an optimization at every step; awkward | Natural — the policy outputs distribution parameters |
| Policy type | Deterministic (from the argmax), unless post-hoc randomized | Naturally stochastic — useful for exploration and for games with no deterministic optimum (e.g. rock-paper-scissors) |
| Sample reuse | Off-policy variants (Q-learning) reuse old data via replay | Classic REINFORCE/A2C are on-policy — each batch of data is used once, then discarded |
| Convergence behavior | Can be unstable (deadly triad) but sample-efficient when stable | More stable gradient-following, typically less sample-efficient |
By the end of this subject you should be able to derive the REINFORCE gradient estimator, explain in one sentence why subtracting a baseline does not bias it, define the advantage function and say why it is the natural quantity a critic should estimate, and sketch the A2C training loop including the design choice between a shared and a separate actor/critic network.