Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Compute Eligibility Traces and Backward-View Updates by Hand

A robot vacuum's episode visits rooms in this order: Kitchen, Hallway, Kitchen, Living Room, over 4 time steps, with \gamma = 0.8 and \lambda = 0.5 (so the per-step trace decay factor is \gamma\lambda = 0.4). Use accumulating eligibility traces, e_t(s) = \gamma\lambda\, e_{t-1}(s) + \mathbb{1}[s_t=s], starting from e_{-1}(s)=0 for all states.

  1. Compute the eligibility trace value for Kitchen, Hallway, and Living Room immediately after each of the 4 steps (build the full table).
  2. At step 4 (visiting Living Room), suppose the observed TD error is \delta_4 = 2.0 and the learning rate is \alpha = 0.1. Compute the update applied to V(\text{Kitchen}), V(\text{Hallway}), and V(\text{Living Room}) at this step.
  3. Explain in one or two sentences why Kitchen receives a larger update than Hallway despite both being visited earlier than the current step, and why this is the desired behavior for a delayed-reward problem.

Share this question

← Back to Reward Design & Delayed Credit Assignment practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.