Design the MDP for a Notification Send/Wait Agent
You are building an RL-based notification agent that decides, once per
idle period, whether to send or wait. The product goal is to
maximize long-term engagement without driving users to mute the
channel.
- Propose a state representation. What information must it contain for the process to be (approximately) Markov, and what happens to the Bellman equations if you leave out the user's fatigue history?
- Propose a reward function for
sendandwaitthat captures the trade-off between short-term opens and long-term mute risk. - Explain, in your own words, why this problem cannot be solved as a standard supervised classification problem ("predict P(open)") even though a supervised click model could be a useful component of your system.
1. State representation and the Markov property
A reasonable state includes: notifications sent in the last N days, time since the last notification, time since the last open, a rolling open rate over the last K notifications, user lifecycle stage, and time-of-day/day-of-week context. The Markov property is an assumption about the state, not a fact about the user: it holds only to the extent the state captures everything relevant to predicting the future. If fatigue history is left out, two users who look identical on the remaining features but have very different recent send counts will be treated identically by the policy — a user who has been hammered with notifications and a fresh user get the same action. The Bellman equations are still written the same way, but they stop being exactly true for the process, because the same "state" now hides two different underlying dynamics; value estimates become systematically biased (averaging over both kinds of user) and the resulting policy will over-send to fatigued users. The fix is always to enrich the state (add a fatigue/recency feature), not to abandon the MDP framing.
2. Reward function
A workable per-step reward: +v_open if the notification is opened,
a small negative -c_fatigue on every send regardless of outcome
(an attention tax), and a large negative -v_mute if the action
causes the user to mute the channel (an absorbing, catastrophic
outcome), e.g. r = 1[opened]*1.0 - 1[sent]*0.05 - 1[muted]*20.
wait gets a small or zero reward (perhaps a tiny opportunity cost
for delaying) and no fatigue tax. The asymmetry — mute cost far
exceeding a single open's value — is what makes the optimal policy
back off as a user approaches the mute threshold, rather than a
hand-tuned "send at most 3x/week" rule.
3. Why not supervised classification
A supervised P(open) model answers a single-step question in
isolation: "will this notification be opened?" It has no mechanism
to account for the fact that sending now changes the state the
agent will face next (higher fatigue, closer to the mute threshold),
and therefore changes the value of future actions. Maximizing P(open)
greedily at every step is exactly the myopic policy this subject
warns about — it chases immediate reward and ignores delayed cost,
producing a "send-happy" policy that drives up mute rate. The RL
framing is what lets the agent trade off a slightly lower open
probability now against a much lower mute probability later, because
it optimizes the discounted return, not the immediate reward. That
said, a supervised P(open) model is genuinely useful as an input
feature or as a component of the reward/transition model — RL doesn't
replace supervised learning here, it wraps it in a longer-horizon
objective.
Share this question