Intermediate
Open
Pro
Solve a Small Bellman System by Hand
Consider a 2-state simplification of the notification MDP: states
OK and Fatigued, absorbing Churned state with value 0. Under the
fixed policy "always send," \gamma = 0.8:
| From | Reward | To OK |
To Fatigued |
To Churned |
|---|---|---|---|---|
OK |
+0.6 | 0.7 | 0.3 | 0.0 |
Fatigued |
−0.3 | 0.2 | 0.6 | 0.2 |
- Write the Bellman expectation equations for V(\text{OK}) and V(\text{Fatigued}) and solve the resulting linear system.
- Suppose instead the agent could
waitinFatigued, with reward -0.05 and transitions \{OK: 0.5,Fatigued: 0.4,Churned: 0.1\}. Using your V(\text{OK}) from part 1, compute Q(\text{Fatigued}, \text{wait}) and compare it to Q(\text{Fatigued}, \text{send}) implied by part 1. What does policy improvement say the agent should do inFatigued? - State, without recomputing, whether this new value can ever be lower than under the "always send" policy, and why.
Share this question