To give context on this video for anyone who doesn't understand. In this video PI* is referring to an idealized policy of actions that result in the maximum possible reward. (In Reinforcement Learning PI is just the actions you take in a situation). To use chess as an example, if you were to play the perfect move at every turn, that would be PI
. Q is some function that tells you optimally, with perfect information the value of any move you could make. (Just like how stock-fish can tell you how many points a move in chess is worth.)
Now my personal comment: for games that are deterministic, there is no difference between a policy that takes the optimal move given only the current state of the board, and a policy that takes the optimal move given even more information, say a stack of future possible turns, etc.
However, in real life, you need to predict the future states, and sum across the best action taken at each future state as well. Unrealistic in the real world where the space of actions is infinite, and the universe to observe is not all simultaneously knowable. (hidden information)
Given the traditional educational background of the professionals in RL, maybe they were referring to the Q* from traditional rl. But I don't see why that would be novel, or notable, as it is a very old idea. Old Old math. From the 60s I think. So I sort of assumed its not. Could be relevant, or just a name collision.