Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — Case Study: Notification Timing & Content Optimization (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Advanced Open Free

Diagnose a Reward-Hacked Notification Policy Permalink →

A messaging app's new RL-based notification policy has been live for six weeks. Dashboards show: daily notification sends per user up 40%, open rate up 12%, and 2-hour engagement (any app open within 2 hours of a send) up 18%. The product team is ready to declare victory. You pull two additional numbers: notification-permission revocations are up 35% over the same period, and 30-day retention for users enrolled in the new policy 6 weeks ago is down 3% relative to a frozen-baseline holdout that has existed since launch.

  1. Explain how a reward function of r = 1{open} - 0.1 * 1{dismiss} (no unsubscribe term) could produce exactly this pattern of dashboard numbers.
  2. Propose a revised reward that would make this failure mode much harder to reach, and explain the mechanism by which each new term helps.
  3. Even with a revised reward, what would you add to the launch process (not the reward function) so this six-week delay before detection doesn't recur?

Share this question

Advanced Open Pro

When to Graduate from a Contextual Bandit to Full RL

Unlock this question →
Advanced Open Pro

An Off-Policy Evaluation Estimate With a Wide Confidence Interval

Unlock this question →
Advanced Open Pro

Design the Guardrail Layer for the Notification Policy

Unlock this question →
Advanced Open Pro

Design the Launch Experiment for the RL Notification Policy

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.