Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Solve a Small Bellman System by Hand

Consider a 2-state simplification of the notification MDP: states OK and Fatigued, absorbing Churned state with value 0. Under the fixed policy "always send," \gamma = 0.8:

From Reward To OK To Fatigued To Churned
OK +0.6 0.7 0.3 0.0
Fatigued −0.3 0.2 0.6 0.2
  1. Write the Bellman expectation equations for V(\text{OK}) and V(\text{Fatigued}) and solve the resulting linear system.
  2. Suppose instead the agent could wait in Fatigued, with reward -0.05 and transitions \{OK: 0.5, Fatigued: 0.4, Churned: 0.1\}. Using your V(\text{OK}) from part 1, compute Q(\text{Fatigued}, \text{wait}) and compare it to Q(\text{Fatigued}, \text{send}) implied by part 1. What does policy improvement say the agent should do in Fatigued?
  3. State, without recomputing, whether this new value can ever be lower than under the "always send" policy, and why.

Share this question

← Back to RL Foundations: MDPs & Value Functions practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.