Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Policy Gradients & Actor-Critic (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Compute a REINFORCE Update With and Without a Baseline Permalink →

A continuous-bidding agent runs a two-step episode with \gamma = 1. State s_0: action sampled with \pi_\theta(a_0 \mid s_0), reward r_0 = 5. State s_1: action sampled with \pi_\theta(a_1 \mid s_1), reward r_1 = -9. Assume you separately know V^\pi(s_0) = 3 and V^\pi(s_1) = -2 (the average return the policy would expect from each state).

  1. Compute the Monte Carlo returns G_0 and G_1, and write the raw (baseline-free) REINFORCE update coefficients for the two timesteps.
  2. Recompute the update coefficients using V^\pi(s_t) as a baseline. What changed, qualitatively, about the signal for the action at s_0?
  3. Prove, in general (not just for these numbers), that adding this baseline does not change the expected value of the gradient estimator.

Share this question

Advanced Open Pro

Choose an Advantage Estimator Under a Bias/Variance Constraint

Unlock this question →
Advanced Open Pro

Debug a Collapsing Shared-Trunk Actor-Critic

Unlock this question →
Advanced Open Pro

Evaluate a Proposal to Add a Replay Buffer to A2C

Unlock this question →
Advanced Open Pro

Design a Policy Head for a Continuous Pricing Action

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.