Can anyone elaborate what “search” means in this context? It seems they are trying to determine the most probable Nash equilibrium but not sure what this mean when applied to RL.
There was some discussion of this in the AlphaGo Zero blog post from a while back: https://deepmind.com/blog/article/alphago-zero-starting-scra...
Imagine searching through the entire tree - billions of nodes due to the branching of chess. Now use RL to help decide where to search and branches to prune