Compute a REINFORCE Update With and Without a Baseline
A continuous-bidding agent runs a two-step episode with \gamma = 1. State s_0: action sampled with \pi_\theta(a_0 \mid s_0), reward r_0 = 5. State s_1: action sampled with \pi_\theta(a_1 \mid s_1), reward r_1 = -9. Assume you separately know V^\pi(s_0) = 3 and V^\pi(s_1) = -2 (the average return the policy would expect from each state).
- Compute the Monte Carlo returns G_0 and G_1, and write the raw (baseline-free) REINFORCE update coefficients for the two timesteps.
- Recompute the update coefficients using V^\pi(s_t) as a baseline. What changed, qualitatively, about the signal for the action at s_0?
- Prove, in general (not just for these numbers), that adding this baseline does not change the expected value of the gradient estimator.
1. Raw returns and coefficients
G_1 = r_1 = -9. G_0 = r_0 + r_1 = 5 - 9 = -4.
Raw REINFORCE update coefficients (the scalar multiplying \nabla_\theta \log \pi_\theta(a_t\mid s_t)) are simply G_0 = -4 and G_1 = -9 — both actions get pushed to become less likely, because the episode's cumulative return from each point was negative, even though the first action was followed by a good immediate reward of +5.
2. With the baseline
G_0 - V^\pi(s_0) = -4 - 3 = -7. G_1 - V^\pi(s_1) = -9 - (-2) = -7.
The coefficient for s_0's action becomes noticeably more negative (−7 instead of −4), because V^\pi(s_0) = 3 says states like s_0 are normally good (a return of +3 expected), so getting an actual return of −4 from there is a much bigger relative disappointment than the raw number suggested — the baseline reveals that this particular action underperformed its state's expectation by more than the raw return implied. This is the opposite of what a naive reading of "baseline shrinks the signal" might suggest: baselines reduce variance across many samples, they do not uniformly shrink magnitude in any one sample, and can enlarge or flip the informative signal for a specific transition, as happened here.
3. General unbiasedness proof
For any function b(s_t) that depends only on s_t, not on the action taken:
Using \pi_\theta(a\mid s_t)\nabla_\theta \log \pi_\theta(a\mid s_t) = \nabla_\theta \pi_\theta(a\mid s_t):
because action probabilities always sum to 1 regardless of \theta, so their gradient is identically zero. This holds for any choice of b(s_t), which is why the baseline can be tuned purely for variance reduction (e.g. set to V^\pi(s_t)) without ever having to worry about correctness — the term it adds to the gradient estimator has zero expectation no matter what function of state it is.
Share this question