Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Choose a GAE Lambda for Two Different Reward Structures

You maintain a shared PPO training library used by two teams:

  • Team A trains a robot-arm policy with a dense reward given every timestep (small penalty for joint jerk, small reward for progress toward the target).
  • Team B trains a policy for a multi-step negotiation agent with a single sparse reward delivered only at the end of an episode (deal value, or 0 if no deal), 40 steps long.
  1. Given a 3-step trajectory with \gamma = 0.95, rewards r_1, r_2, r_3 = 0, 0, 10, and value estimates V(s_1)=1.0,\ V(s_2)=1.5,\ V(s_3)=2.0,\ V(s_4)=0 (terminal), compute the GAE advantage at t=1 for \lambda = 0 and \lambda = 1. Show the TD residuals.
  2. Explain, using your part-1 numbers, why \lambda = 0 would be a poor choice for Team B's setup early in training.
  3. Recommend a \lambda range for each team and justify it in terms of the bias/variance tradeoff.

Share this question

← Back to PPO & Modern Policy Optimization practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.