Advanced
Open
Pro
Select an Action with UCT and Trace the Search Update
An MCTS search is midway through its simulation budget. At the root node, N(s) = 200 total visits have occurred, and c = 1.0. The root has three children:
| Action | Q(s,a) | N(s,a) |
|---|---|---|
| a_1 | 0.70 | 120 |
| a_2 | 0.65 | 70 |
| a_3 | 0.30 | 10 |
- Compute the UCT score for each action and state which one selection would descend into next.
- Suppose the simulation from that node returns a value of 0.9. Describe what happens during backpropagation to the chosen action's Q and N, and to every ancestor on the path to the root.
- After many more simulations, a_3's Q stays low (around 0.28) while its N grows to 60. Explain qualitatively how its UCT score changes, and why this behavior is exactly what you want from the exploration term.
Share this question