Insofar as ARC is being used as a benchmark for code synthesis it might be somewhat successful but it doesn't seem like people are using code synthesis to solve the puzzles so it's not really clear how much success on ARC is going to advance the state of the art in AI and code synthesis according to a logical specification.
It takes a considerable amount of depth in reasoning to see and reason about the patterns / problems / solutions.
Try doing a few of them by hand to see what I mean.
Simulated worlds are complex enough to hide their own flaws just like LLMs are complex enough to lead us to believe they can reason when most of the time they are pattern matching.
I agree that "equivalent to human intelligence" is not a robust way to define general intelligence, but humans are a general intelligence.
In general tech folks are far too beholden to an instinctual and unscientific idea of intelligence as compared between humans, which mostly uses linguistic ability and surface knowledge as a proxy. This proxy might sometimes be useful in human group decision-making, but it is also how dumb confident people manage to fail upwards, and it works about as well for a computer as it does a rat (though it mismeasures in the opposite direction).
I don't see what this has to do with anything. Intelligence is about learning patterns and generalizing them into algorithmic understanding, where appropriate. The number of dimensions latent in the dataset is ultimately irrelevant. Humans live in a 4D world, or 3D if the holographic principle is true, and we regularly deal with mathematics 27 or more dimensions. LLMs build models with at least hundreds of thousands of dimensions.
https://gcptips.medium.com/a-geometric-perspective-on-large-...
As for generalizing to algorithms, LLMs don't yet do this as well as humans, but they do do it:
https://arxiv.org/abs/2309.02390
Finally, there's no intrinsic reason why an AI that can reliably solve deductive problems like ARC would be limited to two dimensions.