Usually in transformer models[1], for each attention head there are 3 vectors of weights, known as q(uery), k(ey) and v(alue). I was assuming that the q in q* applied to the q vector, so q-learning is training this vector. In a transformer model you don't have an objective function for any sort of state evaluation so it can't be the Q you're thinking of.
If they've done something funky with the q vector that could indeed be a breakthrough since a lot of people feel that we are sort of running out of juice with scaling transformers as-is and need a new architecture to really have a step change in capability. That's pure speculation on my part though.
[1] Here's Vaswani et al, the paper that first set out the transformer architecture https://arxiv.org/abs/1706.03762