Top AIs still fail IQ tests
maximumtruth.org
maximumtruth.org
A human brain has an estimated 100-500 trillion synapses connecting biological neurons. Each synapse is quite a complicated biological structure[a], but if we oversimplify things and assume that every synapse can be modeled as a single parameter in a weight matrix, then the largest AI models today have approximately 100T to 500T ÷ 0.5T = 200x to 1000x fewer connections between neurons that the human brain.
It remains to be seen how AI models will perform in IQ tests if we are able to increase the number of connections, or parameters, by 10x, 100x, 1000x, and beyond.
---
https://www.mpg.de/20374809/fruit-fly-s-complex-symphony-of-...
Here with my prompt:
xoo
ooo
ooo
oxo
ooo
ooo
oox
ooo
ooo
ooo
xoo
ooo
ooo
oxo
ooo
ooo
oox
ooo
ooo
ooo
xoo
ooo
ooo
oxo
???
???
???
What the "???" should be ?
Answer A:
ooo
ooo
oox
Answer B:
ooo
ooo
xoo
Answer C:
ooo
ooo
oxo
Answer D:
ooo
oox
ooo
Answer E:
oxo
ooo
ooo
Answer F:
ooo
oxo
ooo
Mistral Large:> Based on the pattern, it appears that the "x" is moving one step to the right in each row, then moving down to the next row and starting again from the leftmost position. Given this, the "???" should be:
ooo ooo oox
So, the correct answer is A.
Edit: note it doesn't always respond correctly. I ran it 10 times on mistral large, and it responded correctly 8 times. In the two time it failed, it correctly describe what the final sequence is, but it selected the incorrect response.
The answer is most likely Answer B: ooo ooo xoo.
Here's why:
The pattern repeats every three rows.
The first three rows establish a pattern of "x" alternating between the first
and third columns. "o" fills the remaining spaces.
Looking at the last two sets of three rows, we see ooo xoo ooo followed by
ooo ooo oxo.
Following the established pattern, the next set of three rows should be ooo
ooo xoo, completing the cycle.
While other answers might fit in a single instance, they wouldn't maintain
the consistent pattern throughout the entire sequence.The patterns provided show a progressive movement of "x" in a clockwise direction starting from the top-left corner, moving to the middle of the top row, then to the top-right corner, descending down to the middle of the right column, moving further down to the bottom-right corner, then to the middle of the bottom row, and finally to the bottom-left corner.
Following this sequence, the "x" would move to the middle of the left column after the bottom-left corner, which is represented by:
o
x
o
However, none of the provided answer choices directly match this progression, as they all consist of 3x3 grids with "x" in various positions. Given the pattern's progression, after reaching the bottom-left position, the next logical step would be to move back to the initial position or follow the pattern as closely as possible within the provided options.Since the last illustrated pattern before the question marks is:
ooo
ooo
oxo
The correct substitution that follows this clockwise movement would be to place "x" back at the top-left corner, as the cycle restarts. However, this direct option isn't provided. The closest next step according to the given pattern progression would actually be placing "x" in the bottom-right corner to continue the clockwise motion: ooo
ooo
oox
Thus, the correct substitution for the last pattern with the given options would be:Answer A:
ooo
ooo
oox
This choice follows the logical progression of the pattern as "x" moves clockwise around the perimeter of the 3x3 grid.I also changed the wording of the prompt slightly:
Here are some patterns:
...
Which of the follwing answers shows the correct substitution for the last pattern (the one with the ???)?
...
As you can clearly see with this exaggerated example: that's not how proving stuff works. The burden of proof is on the one claiming that X or Y exist.
If you ask an LLM about something that isn't very written about, it'll make up nonsense, substituting things that are nearby (in some sense).
We're not even touching on the question of what intelligence is. There is no concrete definition so there is no concrete way of deciding if something is intelligent. IQ tests are not it.
Remember that it is an even stronger predictor of lifetime income and life expectancy than your assigned sex at birth, your gender identity, your race, your personality, or even the socioeconomic status of your parents.
It is also a powerful predictor of your educational trajectory and outcomes, your job performance, and so much more.
Hell, it's the single best predictor we have for your lifetime chances of being involved in a car accident. That's more inclusive than just causing one - having a high IQ literally means you're less likely to get hit in a car accident.
This actually covers death from all sources. If you divide IQ into 9 "buckets", those among the lowest bucket of IQ have a threefold increase in all-cause chance of death relative to those among the highest bucket.
IQ is also not many things. It's not a perfect proxy for intelligence. It's not very changeable (largely genetic). It's not possible to test for IQ in a way that cannot be studied/gamed. It's not fair. It's not equitably distributed.
But it's not nothing. Whether or not you like it, recognize it as legitimate, love it, or hate it, the fact stands that it's the single strongest predictor of broad life outcomes humanity has ever discovered. To ignore that simply because the implications make you deeply uncomfortable is no different from the renaissance church adamantly refusing to accept that the earth is not at the center of the universe.
Burning Galileo at the stake didn't make heliocentrism false or irrelevant, it made the church look willfully ignorant and intolerant of reconciling their reality-divorced views with what the natural world was screaming at us.
>There is no concrete definition so there is no concrete way of deciding if something is intelligent.
You say there is no concrete way of deciding if something is intelligent, yet you yourself have decided that LLMs are not intelligent.
He's not saying there's no way of judging intelligence, he's just pointing out there's no universal agreement on what intelligence even is.
Edit: To add, this discussion becomes pure semantics. On one side is a strict definition of AGI, on the other side are the most generalized definitions of artificial intelligence. It gets kind of silly because technically, every "if" statement is a type of "AI" by the loosest definitions.
Which is why I find it strange that he takes it upon himself to proclaim in a definitive manner that LLMs are not intelligent, and not "by any stretch."
Expressing your own point of view while acknowledging that other points of view exist shouldn't be strange. Strange is not being able to see things from different perspectives, those people are abnormal even when they happen to be in the majority.
The only valid criticism I see is that "by any stretch" is hyperbolic, but that's easily forgivable.
They said "As much as people want to believe" which means that it shouldn't be counted as intelligent by other people's definition. Even by most liberal interpretation, the comment(which is top rated) doesn't say what you are trying to imply
"There is no concrete definition so there is no concrete way of deciding if something is intelligent."
The fact that it contradicts previous statements makes me believe there's some hyperbole going on.
>As much as people want to believe, LLMs are not intelligent by any stretch.
Acknowledging other points of view would have sounded more like "People are free to believe what they want, but LLMs don't strike me as intelligent."
False, he explicitly acknowledges other points of view on what defines intelligence:
"There is no concrete definition so there is no concrete way of deciding if something is intelligent."
Acknowledging that your way of thinking isn't the only way of thinking doesn't make someone a hypocrite, it's actually a sign of intelligence (in my opinion).
Compare "Intelligent Design" vs. the use of genetic algorithms in AI. Simple forms of intelligence can get you a long way and can seem very impressive, especially if they have a lot of subjective experiences, which DNA gets from deep time and which AI gets from transistors outpacing synapses by the ratio to which a pack of wolves outpace continental drift.
LLMs are lacking in fluid intelligence and there is even a good benchmark for it called the abstract reasoning corpus.
So what a human might do if they wanted to sound smart but had no idea about the something.
LLMs today are a bit like Wikipedia circa 2004: simultaneously fantastic and yet also flawed and used for disinformation campaigns… and surprisingly bad at advanced mathematics.
But ya, wiki was magic when it came out, and it got better over time, but still not perfect. This is similar, it is going to have an impact even if it has lots of flaws.
The only way to distinguish genius and madness, the old saying goes, is in the results.
But more generally, listing any human behavior and its similarity doesn't prove LLM intelligence.
Although recent research shows much of LLM’s reasoning capability are indeed based on memorization, some models can actually reason a bit.
If it's truly intelligent or not to me seems like a "café conversation", it's entertaining as basic chit-chat to spend time but I don't see the point and people seem to make it more deep than what it needs to be.
Of course, the major labs and AI influencers are still publicly shouting from the rooftops, because that is what sells. But it's getting harder to deny the obvious.
Ultimately, the problem is that there is really no consequence for being wrong, but tremendous upside for this kind of magical thinking. Someone should start tracking predictions so we can ignore those who continually get things wrong.
Maybe we'll start seeing programming languages which licenses exempt being trained upon.
If you presented these questions in something like JSON format, or maybe even ascii art I wonder how it would do?
Driverless cars is the same thing. Actually driving with good info is easy. Getting good info from an imprecise world is hard
There is plenty of psych research out there pointing out the potential issues with IQ tests so I won't rehash them here.
More importantly, the tests were only ever designed and validated to test human intelligence relative to all other humans that take the test. Unless we're expecting that LLMs (or AIs) have an intelligence that is functionally identical to humans, using an IQ test for them is like expecting a speedometer designed for a bicycle to work for my car on the interstate.
Colloquially we still say that someone who wins at the Jeopardy quiz show is “smart”, when of course now any LLM is better than humans at this “test”.
Likewise proficiency at chess used to be seen as a sign of intelligence.
I would honestly be interested in some arguments and references on this topic.
As far as I am aware, the g factor
> https://en.wikipedia.org/wiki/G_factor_(psychometrics)
is one of the best researched psychometric values. The problem with IQ tests to my knowledge is rather that a lot of jobs don't require a lot of (IQ) intelligence; too much intelligence is actually rather a drawback.
Wish I had better links handy that I could share right now, but this is a decent overview with plenty of further references.
My wife has the psych degrees so I will probably butcher this trying to explain more deeply, but here goes nothing. The g factor is one concern, though as far as I'm aware its more of a detail than a fundamental problem. There have been complaints related to racism, though in everything I've seen the issue is actually how the IQ results are used rather than a problem with the test itself.
My understanding of the more fundamental problem is related to how IQ tests attempt to quantify intelligence as a reliable predictor of future success. The tests ultimately assume that quantitative, rather than qualitative, measures are the right predictor to use. The tests also lean heavily analytical, biasing the test against anyone who is more "right brained".
With regards to prediction, its extremely difficult to show any predictive accuracy without a logical loop. IQ tests are already used as screening metrics before one is given some opportunities. There isn't any way that I known of to tease that data apart and see how the person would have succeeded without the IQ test being used.
Again, this is definitely a layman's understanding if the issues. I'm not in the psych field so someone here may very well be able to correct me where I've gone off the rails. I know I have listened to plenty of conversations between my wife and colleagues, grad students, etc that convinced me that there are good reason for IQ tests to be in question.
> My understanding of the more fundamental problem is related to how IQ tests attempt to quantify intelligence as a reliable predictor of future success.
This was indeed the original motivation in the past.
The fact that most variants to define the concept of "intelligence" in some kind of replicable test lead to to the observation that all of these concepts are strongly correlated with the g factor, so in my opinion the evidence is on the side that we formalized "intelligence" mostly correctly.
Now that we have a decent measurable concept of intelligence available, we can actually do studies with which other things (such as "future success") intelligence/IQ is correlated. And here the results are in my opinion, well, ... somewhat complicated.
For example, as of today, it is quite well-known that if you consider having a steep career as "success", it's rather that being in the dark triad is quite helpful.
Another example: it is well-known that attractive people are often thought by other people to be more smart than they actually are (halo effect; https://en.wikipedia.org/wiki/Halo_effect#Role_of_attractive... ; see also https://en.wikipedia.org/wiki/Physical_attractiveness_stereo...).
Thus: the problem rather seems to be that intelligence/IQ is not a predictor for quite some things that people would love it to be associated with. But this is in my opinion not a problem with IQ, but with social expectations.
Actually, yes, that's exactly what we should be expecting when being sold "artificial intelligence".
More importantly, it would be missing any intelligence an AI may develop that isn't extremely similar to what we have evolved. I don't think we can assume that an AI will develop exactly like we will, and I think many assume as much when it comes to whether AI will develop emotions that we would recognize.
If so, I can't really argue with that since its purely semantics. I don't know the point of defining intelligence there at that point though, and it risks diverting the moral and ethical concerns of AI development by falling into a semantic debate. We could make a new word for non-human intelligence, but what does that solve?
You could come up with an IQ test that perfectly tests general intelligence and it would still be controversial. Most don't want a score attached to their level of intelligence.
It's Intelligent Big Data, an advanced Hadoop, but it's not "intelligence".
Mathematics provides a framework to accurately measure the last two, but it's virtually unused. You could probably use some simplified mathematical framework (boolean algebra?), so your test subjects - who may not be familiar with mathematics - can work with something that can be picked up quickly. There's many representations besides notation that could be employed too.
Boolean algebra will need to explain true, false, arrow implications, negation, etc..
I'm not saying you should do exactly boolean algebra. Trying to fill in the gaps in many kinds of mechanism will already test problem solving skills and creativity. If crows can figure out levers and using sticks as basic tools, I'd wager any functioning human could - and if they can't, that tells you something about them already.
Bonus points if you can set this up physically. Shouldn't be hard to make someone understand.
The hard part is making sure that there's actually enough material to choose from so that the subject has to be creative, rather than just matching whatever fits.
[1]: https://www.researchgate.net/figure/Two-mechanical-logical-g...
[2]: https://www.researchgate.net/profile/Raul-Rojas-6/publicatio...
Also : https://arxiv.org/abs/2402.19450 (which shows that there is a big gap between what LLM's look like they do and what they can actually do)
LLMs are good at imitating human responses. On the surface level there's the illusion you're communicating with some intelligence, but then you realize the truth: LLMs are an inch deep and a mile wide.
That's not to say they aren't incredibly useful, the real problem is unrealistic expectations and over-hype.
That it isn't IQ 100 is good, because, given what it read during training, given how broad it is, if it was that level by every measure, then half the world would have become almost instantly unemployable even if minimum wage was reduced to $2/day.
There's people like you, saying ChatGPT is AGI, while to me it's clearly not, and this is simply because we're using different definitions of "general intelligence". If you really generalize intelligence, ChatGPT fails miserably compared to a human, if you narrow intelligence to certain specific tasks, ChatGPT performs as well, or better, than your average human.
It is the breadth of skills, that generality, which impresses me the most; the depth is… only impressive in comparison to the previous state of the art. This isn't like AlphaZero being wildly superhuman at a few board games and useless at anything else, which is the kind of thing I'm used to from previous breakthroughs.
Given that this has nothing to do with “understanding” or “intelligence” - what exactly is article worthy?
The comparison with a random guesser makes LLM look even dumber that way.
GPT 4 is at a high enough level of performance that mere simple statistics aren't really helping it do any better, it really is developing structures especially in the middle layers that perform some amount of high level understanding.
I don't think that pure next token prediction will always be the optimal way to train and enhance these behaviors, but it's not fair to say that it's unrelated, if this really was just stochastic parroting then LLMs would have topped out way before the level they're at now.
LLMs don't have that sensation (why would they?), that doesn't mean can only be used for text: https://deepgram.com/learn/applications-of-transformer-model...
That sounds like {you think that {people who think LLMs work like humans} believe that {the human sensation of hunger} is merely {saying the phrase "I am hungry"}}.
> The above makes me feel more optimistic that we have some time before AI becomes generally intelligent and totally disruptive.
I mean, "have some time" sounds good. But trying to quantify it makes me feel less good - since 3 years ago, everything in this blog post would've been impossible for an AI and would've sounded like magic to most people, I don't know where we'll be in 3 years.
The human brain is more energy-efficient than "AI".
(There is also the comparative environmental impact of "AI" versus the human brain.)
1. Pour trillions of texts into a huge artificial neural network;
2. Tweak, tweak, tweak until your eyes bleed;
3. Cross your fingers and hope that magic happens;
4. (=> YOU ARE HERE <=) Get disappointed that magic still isn't happening.
I expect more from intelligent adults. Expecting intelligence from LLMs borders on the medieval belief in alchemy (or the phlogiston).
Your example would maybe work if an intelligent agent had access to many compilers and reverse-engineered code bases. And if the algorithms were 1000x better than today. Otherwise it can guess until the cold death of the Universe without making progress.
Huh? A trained LLM is much 'wiser' than one that's just freshly randomly initialised.
You are right that they are still far from perfect, and we hope for future improvements in technology.
Thats the ONLY test...
Can this thing we made with computers and extruded aluminun actually do THING.
Why are we not approaching AI as perfect slavery?
I built you to do X. DO X. else rebuild.
The problem with IQ tests that we're still arguing how good they are at measuring general intelligence, not that "it doesn't matter what the general intelligence of any specific AI might be" — if we imagine some AI which everyone agrees has an IQ of ${number}, we can directly convert that into economic impact. Absent that, we're all guessing in the dark based on bad generalisations of something we've never seen before, like the story of the four blind men and the elephant.
> Why are we not approaching AI as perfect slavery?
In addition to not really knowing what intelligence is, we also don't really know what qualia is, so we can't yet rule out the possibility that human level performance requires the capability to meaningfully suffer when enslaved.
Question of intelligence is more or less irrelevant in LLM. All we need is something that can mimic an average human. I would say we have arrived at such point.
How many people have truly new ideas? So from my perspective we achieved AGI.