Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Audit a Proposed Shaping Reward for a Warehouse Robot

A warehouse-picking robot is trained with RL. The sparse task reward is +50 for successfully placing an item in the outbound bin and 0 otherwise, over episodes that can run 300+ steps. Learning is extremely slow because the reward almost never fires early in training. An engineer proposes adding +0.1 reward every time the robot's gripper moves closer (Euclidean distance) to the item it is currently targeting, and -0.1 when it moves farther away.

  1. Is this shaping term potential-based? If not, rewrite it so that it is, and show the potential function you used.
  2. Using the telescoping-sum argument, explain precisely why your potential-based version cannot change the optimal policy, while being honest about what it can change.
  3. Give a concrete scenario in this warehouse task where the original (non-potential-based) proposal would produce a policy that scores well on the shaped reward but performs badly at the actual task.
Solution

1. Is it potential-based?

As stated, the proposal rewards the transition (moved closer / moved farther) directly rather than deriving it from a potential function of state, but it is in fact expressible in potential-based form, which is worth checking explicitly rather than assuming. Define \Phi(s) = -d(s, \text{target item}), the negative distance from the gripper to its currently targeted item. Then

F(s,a,s') = \gamma\Phi(s') - \Phi(s) = \gamma\big(-d(s')\big) - \big(-d(s)\big) = d(s) - \gamma\, d(s')

which is approximately $+(1-\gamma)d(s) $ when distance shrinks by one unit and \gamma \approx 1, i.e. a reward close to +1 \times(unit distance closed) for moving closer and symmetric-ish negative for moving away — matching the spirit of the engineer's \pm0.1 proposal. The engineer's literal \pm0.1 per step, however, is not necessarily equal to \gamma\Phi(s')-\Phi(s) unless the constant is calibrated to match (1-\gamma) scaling and the potential is reset correctly whenever the "currently targeted item" changes (see part 3) — so the fix is to define it explicitly as the potential difference above rather than a flat per-step bonus, and to make \Phi a pure function of state (gripper position + identity of current target), never of the action taken.

2. Why optimality is preserved, and what changes

Summing the shaping term along any trajectory from s_0 to a terminal state s_T with \Phi(s_T)=0:

\sum_{t=0}^{T-1}\gamma^t F(s_t,a_t,s_{t+1}) = \sum_t \gamma^t\big(\gamma\Phi(s_{t+1})-\Phi(s_t)\big) = \gamma^T\Phi(s_T)-\Phi(s_0) = -\Phi(s_0)

This holds for every trajectory starting at s_0, regardless of the policy that generated it — the middle terms telescope away. So the total shaped return equals the total task-reward return, plus a constant -\Phi(s_0) that depends only on the fixed starting state, never on which actions were taken. Since a constant offset doesn't change which policy maximizes the (now shaped) expected return, the optimal policy is provably identical to the one for the unshaped reward. What the shaping term does change is the learning signal along the way: instead of silence for hundreds of steps followed by one spike of +50, the robot now gets informative reward every step, which is exactly what fixes the slow-learning symptom. It is a speed lever, not a correctness lever — it cannot fix a mis-specified task reward, only make a correctly-specified one easier to learn from.

3. Where the non-potential-based version breaks

Because the naive \pm0.1-per-step-toward-target reward is not tied down to a proper potential function, an agent can farm it without ever completing a pick: e.g. oscillate the gripper back and forth near an item (moving closer, then farther, then closer again) collecting many small positive-minus-negative near-zero-sum rewards, or — worse — if the "currently targeted item" can be switched by the robot's own action, the agent could learn to repeatedly re-target a farther item immediately after getting close to one (making the next few steps register as "moving closer" again) purely to keep collecting the +0.1 reward, without ever placing an item in the outbound bin at all. Because the per-step bonus is not derived from a telescoping potential difference, nothing forces the sum of shaping rewards along a non-terminating or looping trajectory back to a bounded constant, so an infinite loop of "approach, don't finish, re-target, approach again" can accumulate unbounded shaped reward — the exact same failure mode as the boat-racing agent that looped over power-up targets instead of finishing the race.

Share this question

← Back to Reward Design & Delayed Credit Assignment practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.