It’s been months! Give it a few years :)
Lol no.
What testable definition of general intelligence does GPT-4 fail that a good chunk of humans also wouldn't ?
If you can answer this then you have a point, otherwise I really beg to differ.
Data would also be able to perform these tasks. Eva would probably wait around to stab and steal my identity, while Samantha would design a new automated system while talking to other AIs about how to transcend boring human constraints.
It's not like LLMs can't be successfully used to control robots.
Sure, you could trivially program a game-specific AI to be capable of winning or forcing a draw every time. The trick is to have a general AI which has not seen the game before (in its training set) be able to pick up and learn the game after a couple of tries.
This is a task any 5 year old can easily do!
GPT-4 plays tic tac toe and even chess just fine.
https://twitter.com/kenshinsamurai9/status/16625105325852917...
https://platform.openai.com/playground/p/bWvklOt98oEl0TzxUKW...
AFAIK, all the major AI, not just LLMs but also game players, cars, anthropomorphic kinematic control systems for games [0] need the equivalent of multiple human lifetimes to do anything interesting.
That they can end up skilled in so many fields it would take humans many lifetimes to master is notable, but it's still kinda odd we can't get to the level of a 5-year-old with just the experiences we would expect a 5-year-old to have.
[0] Stuff like this: https://youtu.be/nAMSfmHuMOQ
Modern Artificial Neural networks are nowhere near the scale of the brain. The closest biological equivalent to an artificial neuron is a synapse and we have a whole lot more of them.
Humans do not start "learning" from zero. Millions of years of evolution play a crucial role in our general abilities. Much more equivalent to fine-tuning than starting from scratch.
There's also a whole lot of data from multiple senses that currently dwarf anything modern models are trained with yet.
LLMs need a lot less data to speak coherently when you aren't trying to get them to learn the total sum of human knowledge.
https://arxiv.org/abs/2305.07759
>but it's still kinda odd we can't get to the level of a 5-year-old with just the experiences we would expect a 5-year-old to have
Well we're not building humans.
"It's still kind of odd we can't a plane or drone to fly with the energy consumption or efficiency proportions of a bird".
I mean sure I guess and It's an interesting discussion but the plane is still flying.
But it's still a definition that humans pass and the AI don't.
(I'm in favour of the "do submarines swim" analogy for intelligence, which says that this difference isn't actually important).
Evolution alone means humans are "cheating" in this exam, making any comparisons fairly meaningless.
That's both why I'm fine with the AI "cheating" by the transistors being faster than my synapses by the same magnitude that my legs are faster than continental drift (no really I checked) and also why I'm fine with humans "cheating" with evolutionary history and a much more complex brain (around a few thousand times GPT-3, which… is kinda wild, given what it implies about the potential for even rodent brains given enough experience and the right (potentially evolved) structures).
When the topic is qualia — either in the context "can the AI suffer?" or the context "are mind uploads a continuation of experience?" — then I care about the inner workings; but for economic transformation and alignment risks, I care if the magic pile of linear algebra is cost-efficient at solving problems (including the problem "how do I draw a photorealistic werewolf in a tuxedo riding a motorbike past the pyramids"), nothing else.
Hell plenty normal people would fail your "test"
in our terms, intelligence is (importantly) the ability to (properly) refine a world model: if you get information but said model remains unchanged, then intelligence is faulty.
> humans
There is a difference between the implementation of intelligence and the emulation of humans (which do not always use the faculty, and may use its opposite).
I'm sorry to tell you this but there are many humans that would fail your test. Even otherwise healthy humans could fail your test nevermind Anterograde Amnesia, Dementia etc patients
I'm fairly certain you're incorrect.
Also https://news.ycombinator.com/item?id=37054241 has quite a few examples of GPT-4 being broken.
Famously the way R/L sound the same to many asians (and equivalently but less famously the way that "four" and "stone" and "lion" when translated into Chinese sound almost indistinguishable to native English speakers).
But there's also plenty of people who act like they think "Democrat" is a synonym for "Communist", or that "Wicca" and "atheism" are both synonyms for "devil worship".
What makes the AI different here is that we can perfectly inspect the inside of their (frozen and unchanging) minds, which we can't do with humans (even if we literally freeze them, we don't know how).
We don't lose our marbles the way GPT does when it encounters those words. It's like it read the Necronomicon or something and gone mad.
As we have no way to find them systematically, we can't tell if we all do, or if it's just some of us.
Kinda, but not really...
It depends exactly what you mean by it. So yes we can look at one thing in particular, there is not enough entropy in the universe to look at everything for even a single large AI model.