Match a job Paths Subjects Questions Quizzes Pricing
Intermediate Open Pro

Policy Iteration vs. Value Iteration for a Send-Time Optimizer

Your team has a fully specified MDP for send-time optimization: 500 discrete states (fatigue level × recency bucket × lifecycle stage) and 3 actions (send now, wait 6h, wait 24h), with known transition and reward tables estimated from historical logs.

  1. Explain the difference between how policy iteration and value iteration use the Bellman equations, and why policy iteration's outer loop typically needs fewer iterations while value iteration's each iteration is cheaper.
  2. Your infra team says a full policy evaluation to convergence (many sweeps) is expensive to run at 500 states × 3 actions but not prohibitive. Which algorithm would you recommend here and why?
  3. Six months later, the product adds personalization: state now includes a 32-dimensional embedding of user history, effectively making the state space continuous/enormous. Explain why neither policy iteration nor value iteration is viable anymore, and name the category of methods (without deriving them) that this track covers next to handle it.

Share this question

← Back to RL Foundations: MDPs & Value Functions practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.