Long-Term Value & Delayed Reward Systems
A streaming service's recommendation team ships a model change that raises average session watch-time by 6% in a two-week A/B test. Leadership is thrilled. Four months later, 90-day subscriber retention for the treatment group is measurably worse than control. What happened is not a bug — the model did exactly what it was told to do. It optimized watch-time, watch-time went up, and the thing the business actually needs (people staying subscribed) went down, because the two are not the same objective and nobody had reconciled them before launch. This is not a story about a broken reward function; it's a story about the wrong question being asked of the system in the first place. "How do we assign credit to a reward signal correctly" (the mechanics question) is a different problem from "what should the reward signal even be" (the objective question), and this subject is entirely about the second one.
The boundary, stated explicitly. The companion subject, Reward Design & Delayed Credit Assignment, covers the mechanics: given a reward signal you have already chosen, how do you shape it safely (potential-based shaping) and propagate credit for it across a long horizon (n-step returns, eligibility traces), and how do you recognize when a reward is being gamed (reward hacking, Goodhart's Law). This subject assumes those mechanics exist and asks the question that has to be answered first: what should that reward signal measure at all? Should a recommender optimize predicted watch-time, or 90-day retention? Should a notification system optimize open rate, or long-run engagement net of fatigue? These are product and business decisions with real revenue trade-offs, not just modeling choices, and an interview answer that jumps straight to "we'll use TD(λ)" without first defending what the reward is has skipped the harder half of the problem. In an interview, expect to be pushed on exactly this: "why watch-time and not retention?" is a much more common follow-up than "how do you compute the return."
This subject leans on the MDP and value-function formalism from RL Foundations: MDPs & Value Functions and the bandit/contextual-bandit machinery from Multi-Armed Bandits & Exploration and Contextual Bandits for Personalization — those subjects give you the policy; this one gives you the objective that policy should be trained against. It connects forward to Off-Policy Evaluation & Offline RL (you cannot safely test a bad long-term objective live) and to RL in Production: Safe Exploration & Serving. Two case studies apply everything here end to end: Case Study: Notification Timing Optimization and Case Study: Long-Term Engagement Recommender — the latter deliberately overlaps with this subject because it is this subject's central scenario worked through as a full system design.