Speculating from the name only.
Q* might be name derived from Q-learning and A* search algorithm.
In that case it would be informed best best-first search using reinforcement learning.
Q* might be name derived from Q-learning and A* search algorithm.
In that case it would be informed best best-first search using reinforcement learning.
In Q learning, a learner can take some set of actions to maximize a future reward. In this case, the set of actions at each step is the choice of token and the reward is something like 1 if the user liked the response and 0 if they didn’t. Or since it seems they’re applying this to arithmetic, the goal is some formulation of the solution.
Putting these together, it’s possible that Q* is some better way of decoding. Something built on top of the prior probabilities of GPT.