They're not claiming controlling robot arms with language models is a sensible or efficient thing to do, they're trying it to see what happens.
i am tempted to think using it for robotics will become the standard because humans like to have one size fits all solutions.
Trying is stupid.
Trying is fun!
Actually, you don't really need a dedicated model as long as you have the proper adapter like the projector model for vision inputs you just need another for robotic outputs, after that it is just a matter of having the training data.
Just like how you needed to train machine learning models for specific tasks?
ain't stupid if it works, and it works!!! these had been one of "a research team and 15-30 years" problems.