Practice — Model-Based RL & Planning (5 questions)
Advanced
Open
Free
Decide Between Model-Based and Model-Free RL for Two Different Products Permalink →
You advise two teams in the same quarter:
- Team Chess: building an agent to play a board game with fully known rules, deployed against a fast internal simulator that can run millions of games per day on existing infrastructure.
- Team Notify: building an agent to decide which of several push notifications to send a user next, to maximize 30-day retention, using only logged historical interaction data plus live traffic.
- For each team, recommend a model-based or model-free approach (or a specific blend) and justify it using the accuracy of the model each team could realistically build.
- Team Notify's lead proposes: "let's learn a world model of user behavior and use MCTS to plan a multi-step sequence of notifications that maximizes 30-day retention." Explain concretely what would go wrong with this proposal.
- What would you recommend Team Notify build instead, and why does it avoid the failure mode you identified in part 2?
Share this question
Advanced
Open
Pro
Design a Dyna-Style Architecture for a Costly Real-World Environment
Unlock this question →
Advanced
Open
Pro
Diagnose a Planning Failure Caused by Compounding Model Error
Unlock this question →
Advanced
Open
Pro