Advanced
Open
Pro
Detect and Respond to Proxy Metric Divergence
A video platform's recommender is trained to maximize predicted watch-time. Over four successive monthly model releases, the team observes: watch-time per session rising each release (+3%, +4%, +5%, +6%), while a monthly satisfaction survey score has fallen (-1%, -2%, -4%, -7% relative to baseline) and 60-day retention (measured on the cohort exposed to each release, once matured) has also fallen after release 3 and release 4.
- What does this pattern indicate about the relationship between the proxy and the true objective, and at what point should the team have intervened?
- Design a concrete guardrail-metric monitoring scheme that would have caught this earlier than a full 60-day retention read.
- Propose a change to the training objective itself (not just monitoring) that would address the root cause, and explain specifically what problem it solves that better monitoring alone would not.
Share this question