In my own testing, no frontier model knows how to replicate an original 1990s Super Soaker prototype design, which for the most part, should be almost completely possible with Home Depot parts.
They just don't understand PVC parts, triggers, etc.
In my own testing, no frontier model knows how to replicate an original 1990s Super Soaker prototype design, which for the most part, should be almost completely possible with Home Depot parts.
They just don't understand PVC parts, triggers, etc.
What humans "easily" solve in seconds with raw spatial reasoning LLMs often find easier to solve by invoking A* or a constraint solver.
Might be that text data is particularly bad at teaching that to LLMs. Or that being good at spatial reasoning requires true recurrence, and autoregressive chain of thought is a poor substitute. Or it might be that human brain was optimized by evolution for solving spatial problems in open ended 3D environments for hundreds of millions of years, optimized for language for mere hundreds of thousands of years, and only optimized for writing computer code for a few decades at most.
The current frontier is halfway competent at benign closed 2D work, but still completely fumbles anything remotely close to open ended real world 3D work. It's getting better, but very slowly.
Extrapolating the core theory of LLMs - that we can reverse engineer reasoning through language - does that imply that if we train a bird song LLM to predict next “token” (pitch) of a birdsong, that the LLM could excel in a bird flight simulator?
I think it’s pretty clear that this is a dead end.
Do birds expose locomotion-relevant functions specifically through birdsong?
Do we have enough birdsong data available to start solving the inverse problem?
If "yes" on all, then we might be able to do it.
I expect "no" on most of that, for birds. But humans treat language as an interface to their higher cognitive functions, and stockpile language data. That looks an awful lot like a set of two "yes".
The last open question is: is there enough spatial reasoning reflected in the language data we have?
It's plausible that spatial reasoning is too evolutionary old and too low-level, too far removed from higher cognition, to leak into language heavily. And it's also plausible that existing LLM architecture is uniquely poorly suited to learning spatial reasoning - higher cognitive functions involved in things like writing code or even composing poetry are a better fit for the architecture. And it's plausible that we're underestimating just how complex spatial reasoning truly is - Moravec's paradox strikes again.
We know that LLMs perform poorly and improve slowly on spatial reasoning tasks, but not exactly why. And progress on things like ARC-AGI series shows that they're not completely inept.
I think given the fact that spatial reasoning is nearly universal among species, we can very safely assume that it is “too evolutionary old and too low-level, too far removed from higher cognition, to leak into language heavily”
I think this is pretty apparent. It’s very rare for athletes to talk through their actions in high level detail - I saw the ball coming towards me at a 37 degree phi 23 degree epsilon angle at a speed of approximately 20 mph, I estimated it’s time to arrival would be .45 seconds etc. The eye-hand coordination occurs almost completely outside of what you consider conscious awareness. And it’s not easy to describe that’s why athletic coaching is difficult to do through words alone.
As far as ARC-AGI goes it looks like last years models were scoring <5% against their v2 benchmark: https://arxiv.org/pdf/2505.11831
Frankly I don’t understand why you can’t train a multi-modal LLM on video game frame data. Is that just way too compute intensive to do? What am I missing here? Because I think it’s crazy to think that an LLM could learn to think spatially just from reading… even if they’re reading everything that’s ever been written. I think that about summarizes my position.
The issue with multimodal training is that it doesn't seem to bring a step-change improvement in spatial reasoning either. It helps some, but the gain is surprisingly small compared to the data and compute expended. What it helps with the most is, unsurprisingly, spatial reasoning when using image inputs.
Maybe there are gains we don't know how to extract there.
Overall, LLM performance at spatial tasks is improving, especially on things like puzzles, but that mix of "commonsense + spatial" in the same task still eludes them.
But is that actually spatial reasoning? Or is it effectively image generation? Because there’s a difference. Spatial reasoning implies that you could drop it in a video game, give it rules, and let it run. And it would play the game well. Like a flight simulator. That would be true spatial reasoning because spatial reasoning is not just identifying objects but understanding how they interact with one another in a highly quantitative way.
Seems the smart thing to do is not assume an agent will do the right thing. But to create the scaffold / harness that enforces constraints to steer them towards a good result.
Then you can swap out the really smart model for maybe something cheaper.
Of course, there's also no super soaker engineer jobs to take, so I'm sure training sophisticated models to do well in that area is not a high priority for any firms.
I wonder if a more generic lego-manual like task would be more representative. It kind seems like you're testing for AGI.
For me, it's not knowing whether or not it understands there's a big difference between a ball valve and a button release, or that once you start talking about depressing mechanisms for pressure release, you're activating some sort of signals that are too close to triggers (which, is what you want, after all!) and "triggers" are embedded with a very short distance to "guns" in any well-trained model.