Reward Design & Delayed Credit Assignment
Shape rewards without breaking optimality, and assign credit correctly across long, sparse horizons
Learn potential-based reward shaping and why it provably preserves the optimal policy, how reward hacking and Goodhart's Law wreck real products, and the mechanics of credit assignment over long horizons — n-step returns and eligibility traces (TD(λ), forward and backward views) with worked numeric examples.
Practice questions (5)
-
View →
Audit a Proposed Shaping Reward for a Warehouse Robot
Advanced · Free -
View →
Diagnose and Fix an Outrage-Promoting Recommender
Advanced -
View →
A Notification Agent That Learned to Spam
Advanced -
View →
Choose an n for a Delayed-Reward Trading Agent
Advanced -
View →
Compute Eligibility Traces and Backward-View Updates by Hand
Advanced