Advanced
Open
Pro
Compute Eligibility Traces and Backward-View Updates by Hand
A robot vacuum's episode visits rooms in this order: Kitchen, Hallway, Kitchen, Living Room, over 4 time steps, with \gamma = 0.8 and \lambda = 0.5 (so the per-step trace decay factor is \gamma\lambda = 0.4). Use accumulating eligibility traces, e_t(s) = \gamma\lambda\, e_{t-1}(s) + \mathbb{1}[s_t=s], starting from e_{-1}(s)=0 for all states.
- Compute the eligibility trace value for Kitchen, Hallway, and Living Room immediately after each of the 4 steps (build the full table).
- At step 4 (visiting Living Room), suppose the observed TD error is \delta_4 = 2.0 and the learning rate is \alpha = 0.1. Compute the update applied to V(\text{Kitchen}), V(\text{Hallway}), and V(\text{Living Room}) at this step.
- Explain in one or two sentences why Kitchen receives a larger update than Hallway despite both being visited earlier than the current step, and why this is the desired behavior for a delayed-reward problem.
Share this question