Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Reward Penalty vs. Structural Guardrail for a Hard Notification Cap

A messaging app's RL-based notification policy is trained to maximize 30-day retention. Product leadership sets a hard rule: no user may receive more than 3 notifications in a rolling 24-hour window, no exceptions. The ML team implements this by adding a term to the reward function: reward -= 0.5 * max(0, notifications_sent_today - 3).

Three weeks after launch, an audit finds 4,200 users received 5 or more notifications in a single day, all from the learned policy.

  1. Explain mechanically why the reward-penalty approach allowed this to happen, even though the penalty was clearly intended to discourage it.
  2. Redesign the enforcement so the 3-per-day rule literally cannot be violated by any policy version, current or future. Be specific about where in the request path this lives.
  3. Product now wants a softer rule: "prefer not to send more than 2 per day unless predicted value is unusually high." Should this be implemented the same way as the hard 3-cap? Explain the distinction you'd draw and where you'd implement each.

Share this question

← Back to RL in Production: Safe Exploration & Serving practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.