Model-Based RL & Planning
World models, Monte Carlo Tree Search, and Dyna-style architectures — trading compute for real-world samples
A survey of model-based reinforcement learning: learning a transition/reward model and planning against it, Monte Carlo Tree Search (selection, expansion, simulation, backpropagation, UCT), Dyna-style architectures that mix real and simulated experience, why sample efficiency is the core motivation, and when model-based approaches win (accurate-model domains like games) versus lose (hard-to-model domains like open-ended user behavior).
Practice questions (5)
-
View →
Decide Between Model-Based and Model-Free RL for Two Different Products
Advanced · Free -
View →
Select an Action with UCT and Trace the Search Update
Advanced -
View →
Design a Dyna-Style Architecture for a Costly Real-World Environment
Advanced -
View →
Diagnose a Planning Failure Caused by Compounding Model Error
Advanced -
View →
Choose the Right Model-Based Tool for Three Different Systems
Advanced