Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Free

Design the MDP for a Notification Send/Wait Agent

You are building an RL-based notification agent that decides, once per idle period, whether to send or wait. The product goal is to maximize long-term engagement without driving users to mute the channel.

  1. Propose a state representation. What information must it contain for the process to be (approximately) Markov, and what happens to the Bellman equations if you leave out the user's fatigue history?
  2. Propose a reward function for send and wait that captures the trade-off between short-term opens and long-term mute risk.
  3. Explain, in your own words, why this problem cannot be solved as a standard supervised classification problem ("predict P(open)") even though a supervised click model could be a useful component of your system.
Solution

1. State representation and the Markov property

A reasonable state includes: notifications sent in the last N days, time since the last notification, time since the last open, a rolling open rate over the last K notifications, user lifecycle stage, and time-of-day/day-of-week context. The Markov property is an assumption about the state, not a fact about the user: it holds only to the extent the state captures everything relevant to predicting the future. If fatigue history is left out, two users who look identical on the remaining features but have very different recent send counts will be treated identically by the policy — a user who has been hammered with notifications and a fresh user get the same action. The Bellman equations are still written the same way, but they stop being exactly true for the process, because the same "state" now hides two different underlying dynamics; value estimates become systematically biased (averaging over both kinds of user) and the resulting policy will over-send to fatigued users. The fix is always to enrich the state (add a fatigue/recency feature), not to abandon the MDP framing.

2. Reward function

A workable per-step reward: +v_open if the notification is opened, a small negative -c_fatigue on every send regardless of outcome (an attention tax), and a large negative -v_mute if the action causes the user to mute the channel (an absorbing, catastrophic outcome), e.g. r = 1[opened]*1.0 - 1[sent]*0.05 - 1[muted]*20. wait gets a small or zero reward (perhaps a tiny opportunity cost for delaying) and no fatigue tax. The asymmetry — mute cost far exceeding a single open's value — is what makes the optimal policy back off as a user approaches the mute threshold, rather than a hand-tuned "send at most 3x/week" rule.

3. Why not supervised classification

A supervised P(open) model answers a single-step question in isolation: "will this notification be opened?" It has no mechanism to account for the fact that sending now changes the state the agent will face next (higher fatigue, closer to the mute threshold), and therefore changes the value of future actions. Maximizing P(open) greedily at every step is exactly the myopic policy this subject warns about — it chases immediate reward and ignores delayed cost, producing a "send-happy" policy that drives up mute rate. The RL framing is what lets the agent trade off a slightly lower open probability now against a much lower mute probability later, because it optimizes the discounted return, not the immediate reward. That said, a supervised P(open) model is genuinely useful as an input feature or as a component of the reward/transition model — RL doesn't replace supervised learning here, it wraps it in a longer-horizon objective.

Share this question

← Back to RL Foundations: MDPs & Value Functions practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.