What if AGI is not coming?
mindprison.cc
mindprison.cc
GPT-4 Turbo: 31.0
Claude 3 Opus: 27.3
Mistral Large: 17.7
Mistral Medium: 15.3
Gemini Pro: 14.2
Qwen 1.5 72B Chat: 10.7
Claude 3 Sonnet: 7.6
GPT-3.5 Turbo: 4.2
Mixtral 8x7B Instruct: 4.2
Llama 2 70B Chat: 3.5
Qwen 1.5 14B: 3.1
Nous Hermes 2 Yi 34B: 1.5
Notes: 0-shot. Maximum possible is 100. Partial credit is given if the puzzle is not fully solved. There is only one attempt allowed per puzzle. In contrast, humans players get 4 attempts and a hint when they are one step away from solving a group. Gemini Advanced is not yet available through the API.
What I found interesting is how this benchmark reveals a large capabilities gap between the top, large models and the rest, in contrast to existing over-optimized benchmarks.
Just to make sense of your result, can you show your prompt?
When you say 3 prompts and one attempt, what does that mean?
Also regarding 0-shot, did you give the LLM the instruction that is given to human by the game?
If yes, I would count that as one shot as an example of how to properly solve one example puzzle is given.
``` How to Play
Find groups of four items that share something in common.
Select four items and tap 'Submit' to check if your guess is correct.
Find the groups without making 4 mistakes!
Category ExamplesFISH: Bass, Flounder, Salmon, Trout FIRE ___: Ant, Drill, Island, Opal
Categories will always be more specific than "5-LETTER-WORDS," "NAMES" or "VERBS."
Each puzzle has exactly one solution. Watch out for words that seem to belong to multiple categories!
Each group is assigned a color, which will be revealed as you solve ```
Thanks
And frankly, that’s probably not what our brains are doing.
I think that it is! It's much more likely to me that our brains are doing something big and simple than small and complicated. That's the way that nature tends to work. Fitting low-order million-dimensional polynomials would meet that description.
From the double slit experiment, to particle-wave duality, to the particle zoo of the 70s, to quantum chromodynamics, to asymptotic freedom, to more exotic theories like string theory, etc. tells us the complete opposite. Every major discovery in physics in the past 150 years seems to disagree. Things are extremely weird and complicated when we get extremely tiny. Why would our brains be different?
Richard Feynman put it very well:
"The world is strange, the whole universe is very strange, but see when you look at the details then you find out that the rules are very simple, of the game, the mechanical rules by which you can figure out exactly what's going to happen when the situation is simple. It's again this chess game; if you're in just the corner with only a few pieces involved, you can work out exactly what's going to happen. And you can always do that when there's only a few pieces. And so you know you understand it. And yet, in the real game there's so many pieces you can't figure out what's going to happen.
"There's such a lot in the world, there's so much distance between the fundamental rules and the final phenomena that it's almost unbelievable that the final variety of phenomena can come from such a steady operation of such simple rules... But it is not complicated, it's just a lot of it."
ChatGPT by contrast, consumes roughly a gigawatt hour per day serving its users. Yes, to do this it's handling a colossal amount of queries, but that's all it's doing. Your brain handles everything in your body and consciousness, in ways we don't even fully understand, while also letting you think and communicate and reason as a conscious being with self direction.
Moreover, there is evidence that at least part of our brain's functions may be exactly as the other reply here mentions, weird, subatomic and deeply complex in ways that are difficult to get a clear grip on.
Early computers in late 40s were called 'electronic brains' by the media...
Then, one day, they hear about it in the news, because there's now some hype, or some event that makes it newsworthy. This makes it feel like the breakthrough was instantaneous or steep, but in fact has been in the making for decades or even more.
I have no doubt AGI is coming, but it will be gradual and slow. It will be the accumulation of more advances in everything, including hardware, as well as software. It might even include economic changes.
I find that absolutely wonderful and it worked decently well for me (and possibly you.) Now we have a never seen before technology and society will adapt, that's it. No failure.
No. And what's with the "it is so bad yet you used it"? I am very much allowed and required to denounce a system even if I cannot escape it or if I could have somehow profited from it or chosen to use it.
I very much reached the point where I am despite the educational systems I was exposed to. And a system geared towards memorization and regurgitation of data in textual format where pupils can successfully use a chatbot to avoid doing work is certainly failing its goals of educating the youth.
I would point you to the complaints of American teachers about the reading and mathematics levels of students, if only because that is widely accessible. I did not grow up in the US and the school system in my home country is leagues behind the US.
like even restricted to the domain of strictly computation, I'd say it barely scratches the surface... like even if we ignore computer engineering ("the transistor," "silicon microprocessors", etc.)... foundational tech like "compilers" are more significant.
even restricted to modern applications, GPS is more useful and life-changing.
so, no. It's not the biggest step in tech history.
There were hundreds of competitive, even SOTA LLMs before ChatGPT existed. You're basically just proving the parent comment right in how small of a leap ChatGPT is from t5-flan or BERT.
Curious about how many kids in India and China are relying on it.
I'll accept the prospect of hundreds of millions, generously.
And "relying on it" is a strong phrase. Using it as a curiosity, sure. The ones relying on it seem to keep ending up in the news because of how, well, unreliable it is.
In other words, maybe humans have basically solved the optimization problem for the environment we live in. At this point the only thing to compete on is speed and cost.
I actually think it will come from the other direction. That people will get better at asking questions, because there is an automated tool that will build systems to answer larger problems than a single person could quickly answer.
Even if our AI systems have only a minute fraction of von Neumann's intellect, we still have no idea what tomorrow will be like. I'm terrified and excited.
This has changed drastically and thus our definition of smart has too.
If we had more intelligences around to compare with I think we'd find that some are "more intelligent" in that they have all of our capabilities, plus some. And that others are "less intelligent" in that we have all of the capabilities that they have, plus some. And then there would be the "differently intelligent" which have at least one capability that we don't and which lack at least one capability that we have.
Under this lens, I don't know if there's much utility in fine grained comparisons of intelligence re meters and millimeters. The space is discrete: subsets, not metrics.
I don't know if you could ever prove something like this (or maybe we just lack that capability). It seems more like an axiom-selecting notion than something to be argued. Anyhow, it's what my gut says.
If you can find none, is it not the proof that our intelligence is general?
There's also problems of self reference. A Turing machine may be able to solve the halting problem for pushdown automata, but it can't solve the halting problem for Turing machines. Whether or not we're as capable as Turing machines, there's a halting problem for us and we can't solve it.
I'm restricted to mathy spaces here because how else would you construct a well defined question that you cannot answer? But I see no reason why there wouldn't be other perspectives that we're incapable of accessing, it's just that in these cases the ability to construct the question is just as out of reach as the ability to answer it.
You may have heard talk about known unknowns and unknown unknowns, but there are also known unknowables and unknown unknowables, and maybe even unknowable unknowables (I go into this in greater detail here: https://github.com/MatrixManAtYrService/Righting/blob/master...).
In any case, I don't think it's ever valid to take one's inability to find examples as proof of something unless you can also prove that the search was exhaustive.
Instead of AGI, we should call it AHI: artificial human intelligence, or SHI: super human intelligence. That would be much clearer and would sidestep the generality issue.
Also, no one person invented the calculator. The calculator is the culmination of hundreds or even thousands of years of invention. It’s not like the knowledge or creativity is in each of our brains and we could each build a calculator given the requisite materials. It took thousands of lifetimes of ingenuity. So there’s another answer to your question of things we aren’t efficient at solving: building a calculator.
So a bit of a cop-out not wanting to say it outright :)
Why tell the children Santa's not real, when he's the reason they're being so good?
One main reason is that I think people underestimate how much work OUR brains are doing when we interact with LLMs. It seems like the initial "wow" has worn off for many people, but definitely not everybody.
For coding, people will get stuck in loops, trying to get LLMs to modify LLM-generated code
And I think the market will cool down, which seems inevitable considering Nvidia's stock price (I'm a shareholder), and the fact that they seem to be the only ones really making money
If you compare Google after 8 years (2004) to OpenAI after 8 years (2023), the business is uh very different
I’m not invested in any sense in the space. I’m actually more frequently turning off Copilot in VSCode recently. I’d like to see further breakthroughs as much as anyone, but am not holding my breath. In fact, shorting NVIDIA seems like one of the better ideas currently.
But if you believe strongly, shorting AI-enhanced stocks is a great way to capitalize on your prediction. You could also use the short as a hedge for your expectations (either way you win something).
I have no data or sources to back this up.
If we can achieve AGI simply through more and more computation, no matter how novel it is, its ultimately ifs, loops and arithmetic... then surely the human experience is ultimately just a 'wet LLM' (or whatever we end up calling the machine learning technology behind AGI).
Is a soul made out of matrix multiplications and dot products worse than a soul made of neurons?
So that's why it is similarly exclusionary: I think we should try, but also look to see if we can learn more to maybe rule it out altogether.
Who's stopping you?
It's also going to take a while to learn to use the new toys we already have.
We emulate neurons mathematically but it is possible to build efficient analog circuits that emulate them physically.
I doubt any biological system can ever learn all the information gpt4 has learned. The gpt4 learning may not have been power efficient but neither were the first airplanes compared to birds yet today biological flight is rather limited compared to flight that uses technology.
Ever listen to Geoffrey Hinton speak about back propagation vs what biological systems use? Do you think he is wrong about this?
The fact that human brains can become competent at language with training datasets with a size of a tiny fraction of common crawl points to how much more efficient they are
I think you know the difference.
Translating a language isn’t quite the same as understanding it, but this is still impressively data-efficient.
[1] https://twitter.com/hahahahohohe/status/1765088860592394250
"The author has later responded and apologized a flawed / biased methodology which led them to believe that Opus has no prior knowledge of Circassian (i.e. wasn't trained on it)
Apparently it was, and is able to speak Circassian just not perfectly"
https://www.nyu.edu/about/news-publications/news/2024/februa...
How do you provoke a model into being wacky, challenging, and innovative?
Are those the typical qualities of the educated…?
It’s the difference between bluffing and sincerity, or dishonesty and truthfulness. Current LLMs are confident liars.
> we’ll accelerate not a slide into the singularity but a slide into inane banality
Half joking response: Have you looked at all the SEO garbage than any Google search produces these days? We are already in the great age of "inane banality".Worth noting it's been said before for each version of GPT, only to be proven wrong.
But in practice training data is curated and synthetic generated (curated) training data is even better than human data. State of the art LLMs like Phi;2 or the recent GPT-4 killer Claude 3 are trained entirely or mostly on generated data.
The level of results does not scale well down to mass consumer hardware.
And yes I know people can buy an NVIDIA GPU and run these models, but the phone like you said is the most common computer and where this will be hardest to scale too.
It’s why I’m bearish on AI, and I think the pop will be due to being unable to scale down sufficiently
For now. I'm pretty bearish on all things "AI" but of all things one can say about the future, today's hardware is yesterday's news. And in this case, I'd say the same goes for algorithms.
I just don’t think it will live up to the hype people have.
People are seeing Sora and StableDiffusion when they think of AI.
And yes I can run SD on my iPhone, but it’s a poor experience that’s difficult to productize.
The first real products will be so underwhelming compared to what people expect.
Eventually the hardware and algorithms will improve and meet in the middle, but I think it’s so far out, that people will have moved on
The expectation of the current bubble is that something will be delivered soon (in the next year or two)
I don’t see mass consumer hardware scaling up quickly enough in that time to have a product that will match the hype that everyone is showing with cloud based tools.
Where have you seen this? And what is the claim that will be delivered?
I'm absolutely certain something will be delivered in the next year or two. It's probably not AGI.
I think that it does, though. A few years ago, running any remotely complex (and decent) generative AI on anything that wasn't a large compute server was out of the question. Nowadays, you can run a very respectable image or text generation algorithm on a middle-of-the-road gaming PC. Local models may not always be as good as the enormous things that companies run, but they're putting up a fight.
This isn't founded on research, but given how much people were able to scale all of it down, I feel like there's going to be more of this "fat" to trim - data that takes up a lot of space but isn't very important to the final result. Add onto that the constant improvement of hardware, and the lines are going to intersect eventually.
I don’t mean to sound combative, but my point is that the disconnect between what customers will get and what they see is very high. Will your metric meaningfully change that aspect?
Here is a quote from research related to this subject:
> Compared to 2012, it now takes 44 times less compute to train a neural network to the level of AlexNet (by contrast, Moore’s Law would yield an 11x cost improvement over this period). Our results suggest that for AI tasks with high levels of recent investment, algorithmic progress has yielded more gains than classical hardware efficiency.
When you apply the principle of charity you can make their claim increasingly vacuous and eventually true. We're still doing optimization - we're still in the same general structure. The thing is, it becomes absurd when you do that. Its not appropriate to take such a premise seriously. It would be like taking seriously the argument that we haven't had any advancement in software engineering since bubble sort since we're still in the regime of trying to sort numbers when we sort numbers.
Its like, okay, sure, we're still sorting numbers, but it doesn't make the wider point it wants to make and its false even under the regime it wants to make the point under.
This isn't even the only issue that makes this premise wrong. For one, AI research in the 80s wasn't centered around neural networks. Hell even if you move forward to the 90s PAIP puts more emphasis on rule systems with programs like Eliza and Student than it does learning from data. So it isn't as if we're in a stagnation without advance; we moved off other techniques to the ones that worked. For another, it tries to narrow down AI research progress myopically to just particular instances of deep learning, but in reality there are a huge number of relevant advances which just don't happen to be in publicly available chat bots but which are already in the literature and force a broadening. These actually matter to LLMs too, because you can take the output of a game solver as conditioning data for an LLM. This was done in the Cicero paper. And the resulting AI has outperformed humans on conversational games as a consequence. So all those advancements are thereby advances relevant to the discussion, yet myopically removed from the context, despite being counterexamples. And in there we find even greater than 44x level algorithmic improvements. In some cases we find algorithmic improvements so great that they might as well be infinite as previous techniques could never work no matter how long they ran and now approximations can be computed practically.
"Planes and Cars today fundamentally use the same technology we had for almost decades, henceforth ...."
The real question to ask is "does AGI matter"
Sort of reminds me of the late 90s super-proto-VR stuff where people thought any day now we'd be jacking into full immersion (tactile, smell and all) virtual reality.
Don't get me wrong, LLM's are useful tools. But ChatGPT aint Neuromancer. Or even Wintermute. It's Clippy after a few years of community college.
I think the answer is quite obviously no. LLMs can recite their training, and recombine it in ways that correlates strongly to how a human might do so. But creating entirely new knowledge, that goes above and beyond recombinations of what is already known, remains entirely outside the domain of LLMs. An LLM trained on slow classical music is not going to create rap. And an LLM trained on rap is not going to create classical music. And those are trivial examples since it's not entirely new, but just taking a concept and using it a slightly different way than 'normal.' Math, by contrast, is literally creating something from nothing.
And this ability to create something from nothing is probably the most key indicator of intelligence. And we've yet to even step foot on the path towards creating software with this ability.
https://plato.stanford.edu/entries/philosophy-mathematics/
I think your intuition is right but I think it’s more because of the embodiment of humans in the world. The machine wouldn’t invent calculus because it wouldn’t be in a physical world that needed to.
11 dimensional vector calculus is a pretty clear example of recogniting generality between 1, 2 and 3 dimensions we can experience, and extending it to an arbitrary number.
While humans, in their mind, do construct things they did not experience, we develop tools to do that based on what we do (even things like axiom of parallelism: we experience "intersections" or lack of them, so we can ask ourselves "what if two parallel lines do intersect?").
Mathematics has really become abstract in the last 2 centuries, and was pretty tied to physical reality up to that point.
And you are required to see 11 dimension to stand for your body experienceism. You are not allowed to extend linear algebra from 3D to arbitrary dimensions.
LLM nowadays solve insane hard science problem, invent new algorithms that you don't see.
Remixing Jazz tokens might never generate classical music, but an evolutionary algorithm using musical notes and basing its cost function on what humans seem to enjoy might still discover something similar or better. A transformer could then generate infinite versions of that discovery.
More interesting is that these tribes also tend to lack the ability to recall even basic quantities of things if it's more than 2. You can give them 5 balls, but they will have extremely poor recollection of the number after just a few moments. There's these countless things humans need to simply invent, from nothing, over and over to advance to the next technological era. It only seems like a natural application of logic in hindsight, because it actually works!
[1] - https://www.sciencedaily.com/releases/2012/02/120221104037.h...
Of corse one in a million that some will fall into a different evolutionary path.
The number on the success side is statistically way too significant so you can't revolve anything with this one.
I am not saying it's not superior (I do believe it is), just that this does not really support it.
Basically, a single human acting on its own during their single lifetime has almost zero chance of inventing mathematics (as evidenced by the tribed you mention, but also examples of children lost in woods and raised by animals).
And yes, it's going to be a slow process. Human intelligence is not about speed. The capability to do simple arithmetic rapidly correlates with intelligence in humans, yet a calculator can perform said calculations many millions of times faster than the fastest human. Yet that of course doesn't mean said calculator is thus millions of times more intelligent. The question is how long would it take an LLM to discover mathematics starting from the same empty baseline? And, using current methods, the answer is quite simple: never.
https://earthsky.org/earth/fish-can-do-math-cichlids-stingra...
To see how silly these things are imagine somebody claiming to have taught mathematics to one of these numberless tribes by training them to pick a larger group of items when it's blue, and pick a smaller group of items when it's yellow. Somebody claiming that is teaching them math would obviously be looked at like an idiot. It's a circus trick - and a rather apt reflection of what passes for research now a days.
Infinite recursion has to be implemented by iteration in the physical universe.
And then just feed it endless continuous video of the world from a persons perspective (ie, not just random jumps around scenes) How is that any different to the way humans developed our own theories and understanding on the universe we found ourselves in? Obviously a video feed isn't quite the same as having a myriad of senses like we have for interfacing with the world. But theres no reason to say newtons laws need touch or smell to be figured out (maybe that is a bad example for not requiring touch, being about mass and force, but you get the point). Learning language and understanding of the world from scratch might depend on having physcial interaction with the world, sure, but thats again just different sets of input training data that a multimodal model can make connections between if there were pressure to do so.
Then you simply prompt the LLM to define the laws it has observed in a way that can use language to transmit the idea. In the human equivalent, a "prompt" doesn't need to be (but can be) a direct question being asked. It could just be our urge to share ideas, an emotion perhaps that moves us to action. I don't want to go too far on this thought as it is pure speculation without serious scientific backing, but for the sake of the thought experiment with all this in mind you could even say that brains are much like an LLM that is in a constant state of training & updating, but also being "prompted" from our environment and biological feedback loops being fed in. I don't feel like this idea is even that controversial or remotely original.
The difference is that models of today are already trained on the corpus of human knowledge so they already know all of this instead of having to figure it out. But I don't really see an argument for why an LLM wouldn't be able to crystallise observed patterns into mathematical formulae if it had the appropriate influences that would encourage it to do so. Its just that we don't care to train them inefficiently. I imagine this will be tried in the near future though.
I'd love to hear it ;-)
Most important, authors don't know, that all modern AI based on Back Propagation calculations, just because they are easier to implement on cheap old hardware, but natural neurons working on Forward Propagation, which is magnitudes faster on inference.
Unfortunately, for FP we need other hardware, but it is not mean "reaching the limits of hardware scaling", it is just scaling limits for CURRENT hardware, totally other sense.
Sure, if people will play blind and avoid to see obvious things, we will have new AI winter, before somebody will reconsider FP technology.
That's not really true, though. Neural network based approaches are funded, and among those mostly transformers and large language models. Real alternatives aren't funded that much, imo.
I would propose a definition of AGI. "A model capable of effecting the physical world through speech or physical action in a manner indistinguishable from a human."
"Oh don't worry, AGI is coming soon and we'll solve that later" - AI founders
Yet they don't even know how long that is since no-one knows or it never happens. Mistakes in AI are costly and are very expensive.
What if their startup fails before the time arrives because they still cannot make any money and need to constantly raise VC money every week or quarter?
Again, there will only be 90% - 95% of these 'AI' companies that will fail with the 5% to 10% still around including the incumbents.
All humans have the capacity to genuinely learn, create, and think, regardless of how their output appears to you in some subset of interactions with them.
"Some humans sometimes have trouble with critical thinking, or just regurgitate previously-memorized facts" is not in any way equivalent to "LLMs, by their fundamental nature, only have the capacity to produce various recombinations of their training data."
What is happening is the belief that the laws of thermodynamics are probabilistic, like a law that can be broken. Laws like gravity and thermodynamics are deterministic and the hubris of those who make claims of real intelligence in machines we create are going to be as disappointed as those who design perpetual motion machines.
What's lacking is agency/autonomy. I have a bad feeling even 'general autonomy' will take a fraction of the power we're already using which means 'super autonomy'... is probably already possible.
Which means ASI soonish.. which leads to uncontrolled ASI either deliberately or accidently.. which means.. well it's out of our hands at that point. Anything can happen.
This is nonsense statement in and of itself. Its like wondering why an orange fails to turn into a chicken.
There are SO many missing pieces an LLM just doesnt have. LLMs could certainly be a small part of some sort of AGI _system_, but they themselves can never be AGI
Nothing of value will be lost.
>We already have machines that are generally intelligent and require as much energy as a few coalgas lights. Why wouldn’t it be eventually possible to replicate them in brass and steam?
My understanding of the argument of this article is that the conceptual design that replicates intelligence is what the industry has failed to generate today. Simply stating that it might be possible to create is failing to engage; the massive increase in compute power from Babbage to McCarthy didn't give us AGI, because they didn't figure out the right design to reason about anything in even a hundred times a human's energy consumption. If from McCarthy to today we still haven't actually found the proper recipe it just might be worth considering the point the article's making in spite of our other advancements since then.
Probably shouldn’t have mentioned the silicon though.
I mean if you simulate all the chemical processes in some being then yeah probably you get close (assuming you get everything right), let’s assume there is no unknown in how the atoms in our bodies interact.
Sounds like an expensive project.
If you are _not_ talking about a perfect emulation/simulation of such a machine, I will pose your question back to you: why _would_ it be possible to do something we don’t know what it is? Seems rather contrived to say “we can do this thing we don’t know what it is”.
And with the current models, the only thing I see is : "let's add more neurons, give it more data to train, and let's hope a mind will come out of these neurons". It helps understand the human brain, and help humanity and all, but I'm sure we are missing some 'theorical ingredients' to get a recipe for a full working AGI.