Advanced
Open
Pro
Choose a GAE Lambda for Two Different Reward Structures
You maintain a shared PPO training library used by two teams:
- Team A trains a robot-arm policy with a dense reward given every timestep (small penalty for joint jerk, small reward for progress toward the target).
- Team B trains a policy for a multi-step negotiation agent with a single sparse reward delivered only at the end of an episode (deal value, or 0 if no deal), 40 steps long.
- Given a 3-step trajectory with \gamma = 0.95, rewards r_1, r_2, r_3 = 0, 0, 10, and value estimates V(s_1)=1.0,\ V(s_2)=1.5,\ V(s_3)=2.0,\ V(s_4)=0 (terminal), compute the GAE advantage at t=1 for \lambda = 0 and \lambda = 1. Show the TD residuals.
- Explain, using your part-1 numbers, why \lambda = 0 would be a poor choice for Team B's setup early in training.
- Recommend a \lambda range for each team and justify it in terms of the bias/variance tradeoff.
Share this question