Intermediate
Open
Pro
Policy Iteration vs. Value Iteration for a Send-Time Optimizer
Your team has a fully specified MDP for send-time optimization: 500 discrete states (fatigue level × recency bucket × lifecycle stage) and 3 actions (send now, wait 6h, wait 24h), with known transition and reward tables estimated from historical logs.
- Explain the difference between how policy iteration and value iteration use the Bellman equations, and why policy iteration's outer loop typically needs fewer iterations while value iteration's each iteration is cheaper.
- Your infra team says a full policy evaluation to convergence (many sweeps) is expensive to run at 500 states × 3 actions but not prohibitive. Which algorithm would you recommend here and why?
- Six months later, the product adds personalization: state now includes a 32-dimensional embedding of user history, effectively making the state space continuous/enormous. Explain why neither policy iteration nor value iteration is viable anymore, and name the category of methods (without deriving them) that this track covers next to handle it.
Share this question