Match a job Paths Subjects Questions Quizzes Pricing
Overview Read Practice

Practice — RL Foundations: MDPs & Value Functions (5 questions)

Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

Intermediate Open Free

Design the MDP for a Notification Send/Wait Agent Permalink →

You are building an RL-based notification agent that decides, once per idle period, whether to send or wait. The product goal is to maximize long-term engagement without driving users to mute the channel.

  1. Propose a state representation. What information must it contain for the process to be (approximately) Markov, and what happens to the Bellman equations if you leave out the user's fatigue history?
  2. Propose a reward function for send and wait that captures the trade-off between short-term opens and long-term mute risk.
  3. Explain, in your own words, why this problem cannot be solved as a standard supervised classification problem ("predict P(open)") even though a supervised click model could be a useful component of your system.

Share this question

Intermediate Open Pro

Choosing and Justifying the Discount Factor

Unlock this question →
Intermediate Open Pro

Solve a Small Bellman System by Hand

Unlock this question →
Intermediate Open Pro

Policy Iteration vs. Value Iteration for a Send-Time Optimizer

Unlock this question →
Advanced Open Pro

Diagnose a Non-Markov State in a Live Notification System

Unlock this question →

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.