Practice — Case Study: Notification Timing & Content Optimization (5 questions)
Advanced
Open
Free
Diagnose a Reward-Hacked Notification Policy Permalink →
A messaging app's new RL-based notification policy has been live for six weeks. Dashboards show: daily notification sends per user up 40%, open rate up 12%, and 2-hour engagement (any app open within 2 hours of a send) up 18%. The product team is ready to declare victory. You pull two additional numbers: notification-permission revocations are up 35% over the same period, and 30-day retention for users enrolled in the new policy 6 weeks ago is down 3% relative to a frozen-baseline holdout that has existed since launch.
- Explain how a reward function of
r = 1{open} - 0.1 * 1{dismiss}(no unsubscribe term) could produce exactly this pattern of dashboard numbers. - Propose a revised reward that would make this failure mode much harder to reach, and explain the mechanism by which each new term helps.
- Even with a revised reward, what would you add to the launch process (not the reward function) so this six-week delay before detection doesn't recur?
Share this question
Advanced
Open
Pro
An Off-Policy Evaluation Estimate With a Wide Confidence Interval
Unlock this question →
Advanced
Open
Pro