Here's another technical point:
> They're word-association mechanisms with no embodiment and no way to associate the text vectors they manipulate with real-world phenomena.
This is wrong too. RLVR grounds foundational models in reality.
> They're word-association mechanisms with no embodiment and no way to associate the text vectors they manipulate with real-world phenomena.
This is wrong too. RLVR grounds foundational models in reality.
> no way to associate the text vectors they manipulate with real-world phenomena.
Historically, that LLMs were text-only used to be a major argument for why they "lack access to meaning", see the Stochastic Parrot paper and the Octopus paper that it references. But even the authors of those papers have (grudgingly) conceded that the argument no longer holds due to multimodality.