Match a job Paths Subjects Questions Quizzes Pricing

Model-Based RL & Planning

World models, Monte Carlo Tree Search, and Dyna-style architectures — trading compute for real-world samples

Overview Read

Model-Based RL & Planning

Two teams are both building game-playing agents in the same quarter. Team A is training a policy to play a board game with perfectly known rules — every legal move, every state transition, every terminal outcome is exactly computable in code. Team B is training a policy to decide which push notification to send a user next, to maximize 30-day retention. Team A can, in principle, simulate a million games overnight on a cluster and plan several moves ahead before committing to one, because their simulator is the real game, exactly. Team B has no simulator of "what a human will do next" that comes close to being exact — the best they can do is a statistical model fit to noisy, non-stationary, partially observed logs of past behavior, and every plan built on that model inherits its errors.

This contrast is the entire subject in miniature. Model-based RL learns or is given a model of how the world responds to actions — a transition function and a reward function — and uses that model to plan: to look ahead, simulate hypothetical futures, and choose actions accordingly, instead of relying purely on trial-and-error updates to a value function or policy the way the model-free methods in the Value-Based Methods and Policy Gradients & Actor-Critic subjects do. The promise is sample efficiency — squeezing far more learning out of each real interaction — and the catch is that the whole approach is only as good as the model. This subject stays at survey depth on purpose: the goal is to be able to name and compare world models, MCTS, and Dyna architectures fluently in an interview, know the UCT formula and the Dyna loop well enough to sketch them, and reason crisply about when planning against a model helps and when it quietly makes things worse — not to derive every algorithm in the family from scratch.


Model-Free vs. Model-Based: The Core Distinction

Model-free Model-based
What's learned A value function (Q, V) and/or a policy, directly from experience A model of the environment (\hat{P}(s'\mid s,a), \hat{R}(s,a)), used to plan and/or generate simulated experience
How it improves Trial and error on real (or simulated-as-real) transitions Planning against the model, plus optionally still learning from real transitions
Sample efficiency Typically low — needs many real interactions Typically higher — each real transition can inform many planning steps
Failure mode Slow convergence, high sample cost Model error compounds into bad plans ("model bias")
Examples covered elsewhere in this track Q-learning, DQN (Value-Based Methods); REINFORCE, A2C, PPO (Policy Gradients & Actor-Critic, PPO & Modern Policy Optimization) World models, MCTS, Dyna (this subject)

The distinction is not always clean in practice — modern systems (AlphaZero, MuZero, Dreamer) blend both, using a model for planning and a learned value/policy network trained partly model-free-style — but the conceptual split is what interviewers want you to reach for first: does this system decide what to do by simulating consequences before acting, or by having already learned, from raw experience, what tends to work?


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.