Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Choose an Advantage Estimator Under a Bias/Variance Constraint

A robotics-control team is training an actor-critic agent to control a continuous joint torque. Episodes are long (thousands of steps), rewards are dense but noisy (sensor jitter), and the critic is a fresh network still early in training with visibly poor value estimates.

  1. Explain why using the raw Monte Carlo return minus a state-value baseline (G_t - V_\phi(s_t)) is a poor choice for the advantage estimator in this specific setting, even though it is unbiased.
  2. Explain why using the one-step TD error (r_t + \gamma V_\phi(s_{t+1}) - V_\phi(s_t)) alone is also risky here, even though it has much lower variance.
  3. Propose what you would actually use, and describe how you would adjust it over the course of training as the critic improves.

Share this question

← Back to Policy Gradients & Actor-Critic practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.