Practice — Policy Gradients & Actor-Critic (5 questions)
Advanced
Open
Free
Compute a REINFORCE Update With and Without a Baseline Permalink →
A continuous-bidding agent runs a two-step episode with \gamma = 1. State s_0: action sampled with \pi_\theta(a_0 \mid s_0), reward r_0 = 5. State s_1: action sampled with \pi_\theta(a_1 \mid s_1), reward r_1 = -9. Assume you separately know V^\pi(s_0) = 3 and V^\pi(s_1) = -2 (the average return the policy would expect from each state).
- Compute the Monte Carlo returns G_0 and G_1, and write the raw (baseline-free) REINFORCE update coefficients for the two timesteps.
- Recompute the update coefficients using V^\pi(s_t) as a baseline. What changed, qualitatively, about the signal for the action at s_0?
- Prove, in general (not just for these numbers), that adding this baseline does not change the expected value of the gradient estimator.
Share this question
Advanced
Open
Pro