Vision functions the same way as language when inference is done. It's a stream of tokens.
It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.
It's just easier to measure the "right" answer (and deviation therefrom) on a vision task than a language one due to the underspecified nature of language.