I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.
Take that, Jalapeno!
This is about the same rate you get out of Sol Ultraspeed.
Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not a single user thing anymore, which is exactly what they have in the cooker with Astra.
The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much.
I half expect Boston Dynamics to show something like this off in Q4 or whatever.
Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
Patience pays.
But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.
I am relatively certain we have already squarely been beaten in net efficiency at scale.
Productivity is not the only reason to let these meatbags burn oxygen.