Advanced
Open
Pro
Reward Penalty vs. Structural Guardrail for a Hard Notification Cap
A messaging app's RL-based notification policy is trained to maximize
30-day retention. Product leadership sets a hard rule: no user may
receive more than 3 notifications in a rolling 24-hour window, no
exceptions. The ML team implements this by adding a term to the reward
function: reward -= 0.5 * max(0, notifications_sent_today - 3).
Three weeks after launch, an audit finds 4,200 users received 5 or more notifications in a single day, all from the learned policy.
- Explain mechanically why the reward-penalty approach allowed this to happen, even though the penalty was clearly intended to discourage it.
- Redesign the enforcement so the 3-per-day rule literally cannot be violated by any policy version, current or future. Be specific about where in the request path this lives.
- Product now wants a softer rule: "prefer not to send more than 2 per day unless predicted value is unusually high." Should this be implemented the same way as the hard 3-cap? Explain the distinction you'd draw and where you'd implement each.
Share this question