Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Decide Between Model-Based and Model-Free RL for Two Different Products

You advise two teams in the same quarter:

  • Team Chess: building an agent to play a board game with fully known rules, deployed against a fast internal simulator that can run millions of games per day on existing infrastructure.
  • Team Notify: building an agent to decide which of several push notifications to send a user next, to maximize 30-day retention, using only logged historical interaction data plus live traffic.
  1. For each team, recommend a model-based or model-free approach (or a specific blend) and justify it using the accuracy of the model each team could realistically build.
  2. Team Notify's lead proposes: "let's learn a world model of user behavior and use MCTS to plan a multi-step sequence of notifications that maximizes 30-day retention." Explain concretely what would go wrong with this proposal.
  3. What would you recommend Team Notify build instead, and why does it avoid the failure mode you identified in part 2?
Solution

1. Recommendations

Team Chess should lean heavily model-based: the game's dynamics are known exactly (or trivially learnable — the model is the rules), the simulator is cheap and parallelizable, and MCTS-style search (ideally guided by a learned value/policy network, AlphaZero-style) extracts far more decision quality per simulation than either random rollouts or a purely model-free policy trained on the same number of games. This is close to the textbook case model-based RL is built for.

Team Notify should lean model-free, trained on logged data, likely framed as a contextual bandit or short-horizon RL problem rather than long-horizon planning — see part 3.

2. What goes wrong with the world-model-plus-MCTS proposal

Three compounding problems. First, model accuracy: unlike chess, there is no exact or near-exact description of "what a human will do next" — any learned model of user behavior is a noisy statistical fit to non-stationary, partially observed logs (you see clicks and dwell time, not the mental state producing them). Second, compounding error over a multi-step plan: MCTS descends multiple simulated steps into the future using this model to score candidate notification sequences; each step's prediction error compounds into the next, so a plan several notifications deep can rest on a wildly inaccurate simulated user by the final step, while looking confident and well-scored inside the model. Third, consequence asymmetry: in chess, a bad simulated move loses a simulated game — cheap to discover and discard. Here, a policy trained by planning against a wrong model gets deployed and acts on real users before anyone observes that the model was wrong; the cost of a bad plan is real user attention, real churn, and reputational cost, not a discarded simulated trajectory. Non-stationarity compounds this further — even if the model were accurate today, user behavior and the competitive landscape (other apps, trends) drift, so a model trained on last quarter's logs is planning against an increasingly stale picture of the world.

3. What to build instead

A model-free approach trained directly on logged interaction data (the outcomes teams can actually observe and trust), framed as a contextual bandit or short-horizon problem rather than deep multi-step planning — the notification-selection decision has a short enough horizon and immediate-enough feedback (open/no-open, short-term engagement) that a bandit framing captures most of the real signal without needing to trust a multi-step model of the user. Before any new policy touches real traffic, evaluate it against historical logs using off-policy evaluation techniques, so a bad policy is caught against past data rather than discovered by testing it on real users. This avoids the failure mode in part 2 because it never asks a learned world model to be accurate several steps into the future — it either doesn't plan multi-step at all (bandit framing) or evaluates any policy against real historical outcomes before deployment, rather than against a simulated user the team would have to trust blindly.

Share this question

← Back to Model-Based RL & Planning practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.