Advanced
Open
Pro
Choose an Advantage Estimator Under a Bias/Variance Constraint
A robotics-control team is training an actor-critic agent to control a continuous joint torque. Episodes are long (thousands of steps), rewards are dense but noisy (sensor jitter), and the critic is a fresh network still early in training with visibly poor value estimates.
- Explain why using the raw Monte Carlo return minus a state-value baseline (G_t - V_\phi(s_t)) is a poor choice for the advantage estimator in this specific setting, even though it is unbiased.
- Explain why using the one-step TD error (r_t + \gamma V_\phi(s_{t+1}) - V_\phi(s_t)) alone is also risky here, even though it has much lower variance.
- Propose what you would actually use, and describe how you would adjust it over the course of training as the critic improves.
Share this question