Practice — Reward Design & Delayed Credit Assignment (5 questions)
Advanced
Open
Free
Audit a Proposed Shaping Reward for a Warehouse Robot Permalink →
A warehouse-picking robot is trained with RL. The sparse task reward is +50 for successfully placing an item in the outbound bin and 0 otherwise, over episodes that can run 300+ steps. Learning is extremely slow because the reward almost never fires early in training. An engineer proposes adding +0.1 reward every time the robot's gripper moves closer (Euclidean distance) to the item it is currently targeting, and -0.1 when it moves farther away.
- Is this shaping term potential-based? If not, rewrite it so that it is, and show the potential function you used.
- Using the telescoping-sum argument, explain precisely why your potential-based version cannot change the optimal policy, while being honest about what it can change.
- Give a concrete scenario in this warehouse task where the original (non-potential-based) proposal would produce a policy that scores well on the shaped reward but performs badly at the actual task.
Share this question
Advanced
Open
Pro