Advanced
Open
Pro
Diagnose a Non-Markov State in a Live Notification System
Your notification agent's state is (hour_of_day, day_of_week, rolling_open_rate_30d). After launch, you observe that among users
with identical state values, some respond very differently to a
send action depending on whether they were sent a notification in
the last 2 hours (a feature not currently in the state).
- Explain precisely why this observation means the process is not Markov with respect to the current state definition, tying your answer back to the formal definition of the transition function P(s' \mid s, a).
- What symptom would you expect to see in a learned value function (or in the residual error of a fitted Bellman equation) as a result of this omission, even before you knew the specific missing feature?
- Two engineers propose different fixes: (a) add
hours_since_last_notificationto the state, or (b) keep the state as is but train a separate model per hour-of-day bucket. Evaluate both.
Share this question