It's only a matter of time before we have robots powered by large models pretrained to "predict the next token" across a bunch of different sensory modalities -- sight, sound, smell, touch, taste, etc. in a variety of artificial and natural settings, including social-interaction settings. Learning to read, learning to talk, learning to interact with the physical world, and so on -- all of it could very well be built upon the simple idea of learning to "predict the next token."
We live in interesting times.