Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Diagnose a Non-Markov State in a Live Notification System

Your notification agent's state is (hour_of_day, day_of_week, rolling_open_rate_30d). After launch, you observe that among users with identical state values, some respond very differently to a send action depending on whether they were sent a notification in the last 2 hours (a feature not currently in the state).

  1. Explain precisely why this observation means the process is not Markov with respect to the current state definition, tying your answer back to the formal definition of the transition function P(s' \mid s, a).
  2. What symptom would you expect to see in a learned value function (or in the residual error of a fitted Bellman equation) as a result of this omission, even before you knew the specific missing feature?
  3. Two engineers propose different fixes: (a) add hours_since_last_notification to the state, or (b) keep the state as is but train a separate model per hour-of-day bucket. Evaluate both.

Share this question

← Back to RL Foundations: MDPs & Value Functions practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.