Your entire premise is wrong.
Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.
Your entire premise is wrong.
Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.
So What is LLM is used here for? It is used for mere translation between different robots. So it is mostly symbolic translation.
What I am talking about is to translation LLM inference directly to movements. For example, if you ask an LLM, how do I open the microwave door? It will list the steps. I am talking about a system that can go from "put the thing in the microwave", to action steps, without having to never once demonstrate it physically, and do it just from LLM inference.
In short, the way LLMs used here is not (categorically) the way I was asking about.
https://arxiv.org/pdf/2505.23705
https://www.pi.website/download/pistar06.pdf
https://www.pi.website/download/pi07.pdf
The thing literally has a diffusion "action expert" sit in the same attention system as a pre-trained VLM. And the VLM itself is ALSO trained to generate raw actions as a part of the training recipe (the first paper) - it just doesn't do it at inference time. What the "action expert" does is parallelize the action generation process - based on VLM's internal states.
It's exactly the thing you claimed to be impossible. Described in detail in a paper from 2025. What's your excuse?
You already downgraded your claims from "LLMs are irrelevant to robotics" to a measly "you can't train a useful robotics LLM because there's not enough data". And you say that while looking at an LLM that was pre-trained on all of internet scraped and only then reused for robotics.
Both the pool of robotics-relevant data and the performance of foundation model LLMs grow over time. All the companies that are serious about robotics are serious about scaling up data collection.
I'm not going to claim that this "LLM core" approach is the best approach to AI robotics possible - but if you're betting on it failing outright, you're going to be fighting uphill.
This was the claim from the very beginning. You should have asked why I think what I think, instead of leading with "the entire premise is wrong!"...
Your entire premise was wrong at every point, and now you're trying to wriggle your way out of admitting it.
Prove it!
Read the command line prompt: --task="pick up the red cube"
So it should be something like, "put back this slipped cycle chain back on sprocket"..
Read the command line prompt: --task="pick up the red cube"
There is a gif directly below it.
This is a completely open source model and arm you can replicate yourself.