Decide Between Model-Based and Model-Free RL for Two Different Products
You advise two teams in the same quarter:
- Team Chess: building an agent to play a board game with fully known rules, deployed against a fast internal simulator that can run millions of games per day on existing infrastructure.
- Team Notify: building an agent to decide which of several push notifications to send a user next, to maximize 30-day retention, using only logged historical interaction data plus live traffic.
- For each team, recommend a model-based or model-free approach (or a specific blend) and justify it using the accuracy of the model each team could realistically build.
- Team Notify's lead proposes: "let's learn a world model of user behavior and use MCTS to plan a multi-step sequence of notifications that maximizes 30-day retention." Explain concretely what would go wrong with this proposal.
- What would you recommend Team Notify build instead, and why does it avoid the failure mode you identified in part 2?
1. Recommendations
Team Chess should lean heavily model-based: the game's dynamics are known exactly (or trivially learnable — the model is the rules), the simulator is cheap and parallelizable, and MCTS-style search (ideally guided by a learned value/policy network, AlphaZero-style) extracts far more decision quality per simulation than either random rollouts or a purely model-free policy trained on the same number of games. This is close to the textbook case model-based RL is built for.
Team Notify should lean model-free, trained on logged data, likely framed as a contextual bandit or short-horizon RL problem rather than long-horizon planning — see part 3.
2. What goes wrong with the world-model-plus-MCTS proposal
Three compounding problems. First, model accuracy: unlike chess, there is no exact or near-exact description of "what a human will do next" — any learned model of user behavior is a noisy statistical fit to non-stationary, partially observed logs (you see clicks and dwell time, not the mental state producing them). Second, compounding error over a multi-step plan: MCTS descends multiple simulated steps into the future using this model to score candidate notification sequences; each step's prediction error compounds into the next, so a plan several notifications deep can rest on a wildly inaccurate simulated user by the final step, while looking confident and well-scored inside the model. Third, consequence asymmetry: in chess, a bad simulated move loses a simulated game — cheap to discover and discard. Here, a policy trained by planning against a wrong model gets deployed and acts on real users before anyone observes that the model was wrong; the cost of a bad plan is real user attention, real churn, and reputational cost, not a discarded simulated trajectory. Non-stationarity compounds this further — even if the model were accurate today, user behavior and the competitive landscape (other apps, trends) drift, so a model trained on last quarter's logs is planning against an increasingly stale picture of the world.
3. What to build instead
A model-free approach trained directly on logged interaction data (the outcomes teams can actually observe and trust), framed as a contextual bandit or short-horizon problem rather than deep multi-step planning — the notification-selection decision has a short enough horizon and immediate-enough feedback (open/no-open, short-term engagement) that a bandit framing captures most of the real signal without needing to trust a multi-step model of the user. Before any new policy touches real traffic, evaluate it against historical logs using off-policy evaluation techniques, so a bad policy is caught against past data rather than discovered by testing it on real users. This avoids the failure mode in part 2 because it never asks a learned world model to be accurate several steps into the future — it either doesn't plan multi-step at all (bandit framing) or evaluates any policy against real historical outcomes before deployment, rather than against a simulated user the team would have to trust blindly.
Share this question