Practice — RL Foundations: MDPs & Value Functions (5 questions)
Intermediate
Open
Free
Design the MDP for a Notification Send/Wait Agent Permalink →
You are building an RL-based notification agent that decides, once per
idle period, whether to send or wait. The product goal is to
maximize long-term engagement without driving users to mute the
channel.
- Propose a state representation. What information must it contain for the process to be (approximately) Markov, and what happens to the Bellman equations if you leave out the user's fatigue history?
- Propose a reward function for
sendandwaitthat captures the trade-off between short-term opens and long-term mute risk. - Explain, in your own words, why this problem cannot be solved as a standard supervised classification problem ("predict P(open)") even though a supervised click model could be a useful component of your system.
Share this question
Intermediate
Open
Pro