Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Diagnose a Reward-Hacked Notification Policy

A messaging app's new RL-based notification policy has been live for six weeks. Dashboards show: daily notification sends per user up 40%, open rate up 12%, and 2-hour engagement (any app open within 2 hours of a send) up 18%. The product team is ready to declare victory. You pull two additional numbers: notification-permission revocations are up 35% over the same period, and 30-day retention for users enrolled in the new policy 6 weeks ago is down 3% relative to a frozen-baseline holdout that has existed since launch.

  1. Explain how a reward function of r = 1{open} - 0.1 * 1{dismiss} (no unsubscribe term) could produce exactly this pattern of dashboard numbers.
  2. Propose a revised reward that would make this failure mode much harder to reach, and explain the mechanism by which each new term helps.
  3. Even with a revised reward, what would you add to the launch process (not the reward function) so this six-week delay before detection doesn't recur?
Solution

1. Mechanism

With no unsubscribe penalty, and only a small dismiss penalty, the policy's optimal move under this reward is straightforward: send more. Every additional notification carries a positive expected contribution to the reward (some nonzero chance of an open) at a fixed, small cost per dismiss, so as sends increase, so does raw open count and 2-hour engagement — both metrics the reward is directly correlated with. The 35% revocation increase is the cost the reward function never accounted for: users pushed past their tolerance don't dismiss individual notifications (which the reward mildly penalizes), they revoke permission entirely (which the reward doesn't see at all, because a revoked user simply generates no more logged decisions — the negative outcome is invisible to the training signal). The 30-day retention decline is the true cost showing up on the one metric the policy was never optimized for and that takes weeks to mature, which is exactly why it was the last number anyone looked at.

2. Revised reward

Add a heavily-weighted unsubscribe/revocation penalty, anchored to the estimated value of all future notification-driven engagement lost when a user revokes permission (not an arbitrarily chosen constant), and replace the raw open-rate term with a short-horizon surrogate for retention (e.g., a 7-day session-return delta relative to what would have been predicted anyway) rather than a same-notification open flag. The unsubscribe term directly makes the policy's worst failure mode (permanently losing the channel) enormously costly rather than invisible, closing exactly the gap that let sends climb unchecked. The surrogate-reward swap matters because it decouples "did this specific notification get opened" (easy to farm by sending more) from "did overall engagement over the following week actually improve" (much harder to farm by volume alone, because spamming a user typically depresses, not raises, their engagement over the following week once fatigue sets in).

3. Process fix, independent of the reward

Add a mandatory off-policy evaluation gate before any policy version is allowed past a small capped canary, using the propensities logged at serving time to estimate expected return (including the unsubscribe-risk term) on historical traffic before it ever reaches live users at scale — this catches a reward-hacking policy in hours against logs rather than after six weeks of live damage. Also add unsubscribe/revocation rate as an independently monitored guardrail metric with an automatic circuit breaker (suppress non-critical sends for a user, or halt the experiment) that fires on its own threshold, completely separate from whatever the policy's own reward estimate says — so detection does not depend on someone remembering to check a 30-day-lagging metric before declaring victory on 2-hour numbers.

Share this question

← Back to Case Study: Notification Timing & Content Optimization practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.