Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Compute a REINFORCE Update With and Without a Baseline

A continuous-bidding agent runs a two-step episode with \gamma = 1. State s_0: action sampled with \pi_\theta(a_0 \mid s_0), reward r_0 = 5. State s_1: action sampled with \pi_\theta(a_1 \mid s_1), reward r_1 = -9. Assume you separately know V^\pi(s_0) = 3 and V^\pi(s_1) = -2 (the average return the policy would expect from each state).

  1. Compute the Monte Carlo returns G_0 and G_1, and write the raw (baseline-free) REINFORCE update coefficients for the two timesteps.
  2. Recompute the update coefficients using V^\pi(s_t) as a baseline. What changed, qualitatively, about the signal for the action at s_0?
  3. Prove, in general (not just for these numbers), that adding this baseline does not change the expected value of the gradient estimator.
Solution

1. Raw returns and coefficients

G_1 = r_1 = -9. G_0 = r_0 + r_1 = 5 - 9 = -4.

Raw REINFORCE update coefficients (the scalar multiplying \nabla_\theta \log \pi_\theta(a_t\mid s_t)) are simply G_0 = -4 and G_1 = -9 — both actions get pushed to become less likely, because the episode's cumulative return from each point was negative, even though the first action was followed by a good immediate reward of +5.

2. With the baseline

G_0 - V^\pi(s_0) = -4 - 3 = -7. G_1 - V^\pi(s_1) = -9 - (-2) = -7.

The coefficient for s_0's action becomes noticeably more negative (−7 instead of −4), because V^\pi(s_0) = 3 says states like s_0 are normally good (a return of +3 expected), so getting an actual return of −4 from there is a much bigger relative disappointment than the raw number suggested — the baseline reveals that this particular action underperformed its state's expectation by more than the raw return implied. This is the opposite of what a naive reading of "baseline shrinks the signal" might suggest: baselines reduce variance across many samples, they do not uniformly shrink magnitude in any one sample, and can enlarge or flip the informative signal for a specific transition, as happened here.

3. General unbiasedness proof

For any function b(s_t) that depends only on s_t, not on the action taken:

\mathbb{E}_{a_t \sim \pi_\theta}\big[\nabla_\theta \log \pi_\theta(a_t\mid s_t)\, b(s_t)\big] = b(s_t)\sum_a \pi_\theta(a\mid s_t)\nabla_\theta \log \pi_\theta(a\mid s_t)

Using \pi_\theta(a\mid s_t)\nabla_\theta \log \pi_\theta(a\mid s_t) = \nabla_\theta \pi_\theta(a\mid s_t):

= b(s_t)\sum_a \nabla_\theta \pi_\theta(a\mid s_t) = b(s_t)\,\nabla_\theta\!\left(\sum_a \pi_\theta(a\mid s_t)\right) = b(s_t)\,\nabla_\theta(1) = 0

because action probabilities always sum to 1 regardless of \theta, so their gradient is identically zero. This holds for any choice of b(s_t), which is why the baseline can be tuned purely for variance reduction (e.g. set to V^\pi(s_t)) without ever having to worry about correctness — the term it adds to the gradient estimator has zero expectation no matter what function of state it is.

Share this question

← Back to Policy Gradients & Actor-Critic practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.