But I can ask 100 causally masked questions against common prefix, and get 100 answers, all in a single PP pass using any existing "classical" attention transformer model? Like, I had the impression that is what everyone was doing for classification already?
Is the difference "we did RL to tune logit distribution"? Because I really do not see anything new there. What is the difference?