The surprising effectiveness of test-time training for abstract reasoning [pdf]
mit.edu
mit.edu
We still have a long way to go for the grand prize -- we'll be back next year. Also got some new stuff in the works for 2025.
Watch for the official ARC Prize 2024 paper coming Dec 6. We're going to be overviewing all the new AI reasoning code and approaches open sourced via the competition [3].
[1] https://deepmind.google/discover/blog/ai-solves-imo-problems...
The point of the contest is to measure intelligence in general-purpose AI systems: it does not seem in the spirit of the contest that this AI would completely fail if the test was presented on a hexagonal grid.
Solving ARC-AGI represents a material stepping stone toward AGI. At minimum, solving ARC-AGI would result in a new programming paradigm. It would allow anyone, even those without programming knowledge, to create programs simply by providing a few input-output examples of what they want.
This would dramatically expand who is able to leverage software and automation. Programs could automatically refine themselves when exposed to new data, similar to how humans learn.
If found, a solution to ARC-AGI would be more impactful than the discovery of the Transformer. The solution would open up a new branch of technology.
This program does not represent a "new paradigm" because it requires a bunch of human programming work specifically tailored to the problem, and it cannot be generalized. If software like this wins the contest that really shows the contest has nothing whatsoever to do with AGI.Currently, any claim about AGI other than "we're probably not anywhere close to strong AGI" is simply false.
Of course a lot depends on one's definition of AGI. From another perspective, one could argue that ChatGPT 4 and similar models are already AGI.
The real pressure is the private hold-out set and the variations that can be added to counter this aspect.
A true AGI would be able to solve anything thrown at it which is where the authors are trying to lead AI engineering towards since LLMs have pretty much taken over.
If it starts getting too easy, they just reconsider and add harder problems.
It's like how we don't talk about the Turing Test anymore as it's no longer the best metric to determine real intelligence.
The authors are signalling to the industry that new ideas are needed and the monetary aspect is to show how serious they are about it.
It's good because as per above we have research being thrown at it which means we can iterate until we perhaps find another breakthrough.
Insofar as ARC is being used as a benchmark for code synthesis it might be somewhat successful but it doesn't seem like people are using code synthesis to solve the puzzles so it's not really clear how much success on ARC is going to advance the state of the art in AI and code synthesis according to a logical specification.
I don't see what this has to do with anything. Intelligence is about learning patterns and generalizing them into algorithmic understanding, where appropriate. The number of dimensions latent in the dataset is ultimately irrelevant. Humans live in a 4D world, or 3D if the holographic principle is true, and we regularly deal with mathematics 27 or more dimensions. LLMs build models with at least hundreds of thousands of dimensions.
https://gcptips.medium.com/a-geometric-perspective-on-large-...
As for generalizing to algorithms, LLMs don't yet do this as well as humans, but they do do it:
https://arxiv.org/abs/2309.02390
Finally, there's no intrinsic reason why an AI that can reliably solve deductive problems like ARC would be limited to two dimensions.
It takes a considerable amount of depth in reasoning to see and reason about the patterns / problems / solutions.
Try doing a few of them by hand to see what I mean.
Simulated worlds are complex enough to hide their own flaws just like LLMs are complex enough to lead us to believe they can reason when most of the time they are pattern matching.
In general tech folks are far too beholden to an instinctual and unscientific idea of intelligence as compared between humans, which mostly uses linguistic ability and surface knowledge as a proxy. This proxy might sometimes be useful in human group decision-making, but it is also how dumb confident people manage to fail upwards, and it works about as well for a computer as it does a rat (though it mismeasures in the opposite direction).
I agree that "equivalent to human intelligence" is not a robust way to define general intelligence, but humans are a general intelligence.