That pretty much already exists. Look at DeepMind's Gato: all tasks and modalities are simply sequences of tokens, everything from 'predict English text' to 'predict VAE image token sequences' to 'predict robotic arms commands and movements IRL'.
In any case, I'm not sure Gato qualifies as a "large" model with 1.2B parameters -- it's kinda right below the threshold at which it could or would start exhibiting emergent behaviors. Maybe a new Gato with 10's or 100's of billions of parameters operating in the physical world?
I hope it didn't all get rolled up into Gemini and become a state secret they'll never publish on again, or lost in the shuffle in the chaos of the DeepMind/Brain merger/liquidation.
That's the most likely explanation, in my view.