How do you apply Q-learning to LLMs? They would need to estimate the action-value of each action in a given state. How can you pick the best action when you only have a scoring function that would rate how good each action is, but the action space is infinite. You can't pick the best action.
It is probably something else, like Decision Transformer.