Take the classic trick question, for example: "A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?"
Most people give a wrong answer because they, too, "pattern match".
I asked this question in a college-level class with clickers. For the initial question I told them, "This is a trick question, your first answer might not be right". Still less than 10% of students got the right answer.
Students can keep ahead of it with training in specific fields and it has a few weaknesses in specific skills but I think someone could make a reasonable claim that ChatGPT has superior general intelligence.
[0] https://medium.com/@soltrinox/the-i-q-of-gpt4-is-124-approx-...
It's abhorrent any time we measure pure distilled intelligence.
When asked to come up with any non-basic novel algorithm and data structure, it creates nonsense.
Especially when you ask it to create vector instruction friendly memory layouts and it can't code in its preferred way. I had some fun trying to make it spit out a brute-force-ish solver for a problem involving basic orbital mechanics and some forces. Wouldn't even want try something more complicated. It can do generalized solvers somewhat, since it can copy that homework, but none that can express the kinds of terms you'd be working with (despite those also having code available in some research papers).
Speaking of which, it cannot even figure out some basic truths in orbital mechanics that can be somewhat easily derived from the formulas commonly given, nine times out of ten (you can get there if you're very patient and are able to filter its wrong answers).
But at the end of the day it was still a valuable tool to me as I was learning these things myself, since despite being often wrong, it nevertheless spat out useful things I could plug into Google to find more trustworthy sources that would teach me. Really neat if you're going in blind into a new subject.
Despite me being someone who is generally impressed by the best LLMs, I think this says more about IQ tests than it does any AI.
Which isn't to shame those tests — we made those tests for humans, we were the only intelligence we knew of that used complex abstract language and tools until about InstructGPT — but it does mean we should reconsider what we mean by "intelligence".
My gut feeling is that a better measure is how fast we learn stuff. Computers have a speed advantage because transistors outpace synapses by the degree to which a pack of wolves outpaces continental drift (yes I did do the calculation), so what I mean here is how many examples rather than how many seconds.
But as I say, gut feeling — this isn't a detailed proposal for a new measure, it likely needs a lot of work to even turn this into something that can be a good measure.
https://g.co/gemini/share/94238c5ed174 / https://archive.is/ZTkjI
bat=ball+$1
(ball+$1)+ball=$1.10
2ball + $1 = $1.10
2ball = $0.10
ball = $0.05
This is actually a great example of what I'm talking about. All this talk about how useful AIs are, and how humans have flaws and yet the conclusion is always the same. AIs are only useful for tasks that are relatively easy and have a higher tolerance for failure.
Human intelligence is not benchmarked by its lowest common denominators (just like how we don't judge LLM's on the basis of tiny 100M parameter models).
Your previous reply was to:
> Why would we try to answer that question without discussing how humans think?
And by replying to that with
> Coz we simply don't understand how we understand, period.
Well, when we don't understand how we understand, that is exactly when we should be discussing how humans think. Or at least how we think we think. And the bat and a ball example relates to research about how we think we think.
So yeah, your reply definitely comes across as saying the first half, "So we shouldn't discuss it", while finishing it with "period" also suggests "or try to understand it at all".
The post you are replying to does not suggest that we should. If anything, it is suggesting the opposite - that we should be considering the full range of human abilities (or at least those that are effective in solving complex problems) when addressing the question you have quoted.
I agree with most of what viraptor has said in this thread, but not in this particular case.
5 years ago, it was still common to read things like "AIs will never be as intelligent as mice, let alone humans", today, it's "sure, AIs are as intelligent as some humans, but not as intelligent as the right humans".
Noticed how everyone dropped the Turing Test like a hot potato as the gold standard for intelligence, the moment it became apparent that LLMs were about to pass it? Try to find a recent high-profile article invoking the Turing Test. Crickets. The intellectual dishonesty is nauseating.
The entire discussion is dominated by smart people who are scared shitless that AI is going to show them just how ordinary they are in the grand scheme of things.
I don't believe we train The LLMs for those things specifically. It will be interesting to see if some datasets for this appear. I think we can still make huge improvements by just caring about that more.
We've got Samantha though, so I hope we will see those attributes too. https://erichartford.com/meet-samantha
Ah, wait. That only counts as evidence of intelligence when humans do it, right?
Ah, wait. That only counts as evidence of intelligence when AIs do it, right?
Only if you count humanity as a whole do we beat an LLM at everything.
1. You're counting "our day job" as one task, while counting all individual prompts an AI can answer as each being their own task. This is obviously misleading (and I could just as easily perform the same compression in the opposite direction - chatbots can only do one task: "be a chatbot").
2. You're not controlling for training. It's already meaningless to compare the intelligence of one entity trained to do a task with another entity that was not trained to do that task.
But even ignoring those fallacies, what you've written here is still not true. 99% of things an LLM could supposedly "out-perform" a human at, the human would actually outperform if you provided that human with the same text resources the LLM used to conjure its answer. Regurgitating facts is not evidence of intelligence, and humans can do it easily (and do it without hallucinating, which is key) if you just give them access to the same information.
But when you go in the opposite direction, achieving parity is no longer so simple. When LLM's fail to do math, fail to strategize, fail to process rules of abstract games, etc. there is no textbook or lecture or article you can provide to the LLM that magically makes the problem go away. They are fundamental limitations of the AI's capabilities, rather than just a result of not possessing enough information.
What many people don't seem to realize is, when merely getting up in the morning and brushing your teeth you are already exercising more intelligence than any AI has ever possessed. (Anyone who has ever worked in robotics or visual processing can attest to that enthusiastically.) So don't even get me started on actual critical thinking.
1. This is indeed a simplification, but for any single task in your day job, it would be those tasks where you have the most experience. For example, I used to write video games, AI does a better job of game design than me, but I'm the better programmer.
2. Unimportant, as the consideration I was rejecting was performance in tasks.
As it happens, some of my other recent messages demonstrate that I agree they are low intelligence for this exact reason.
> 99% of things an LLM could supposedly "out-perform" a human at, the human would actually outperform if you provided that human with the same text resources the LLM used to conjure its answer
Could I pass a bar exam of a medical exam, by reading the public internet, with no notes and just from memory, which is what a base model does?
Nope.
Could I do it with a search engine, which is what RAG assisted LLMs do?
Perhaps.
> humans can do it easily (and do it without hallucinating
Hell no we mess that up almost constantly.
> When LLM's fail to do math, fail to strategize, fail to process rules of abstract games, etc. there is no textbook or lecture or article you can provide to the LLM that magically makes the problem go away
I'm sure I've seen this done. I wonder if I'm hallucinating that certainty…
You're -still- ignoring the fact that these models spend millions of GPU-hours in training. I'm sure you could manage.
> Hell no we mess that up almost constantly.
"Almost constantly"? Is this satire? I'd fire any such person, and probably recommend them psychiatric treatment.
AI-hype people really think so little of human beings? I certainly hope my pilot isn't "almost constantly" hallucinating his aviation training.
> I'm sure I've seen this done. I wonder if I'm hallucinating that certainty…
Show me the conversation.
Understanding in this sense seems different from the memorizing+flexible retrieval we know LLMs exceed at because it extends much further beyond its training distribution. If I ask a (good) medical professional a question unlike anything they’ve ever seen before, they’ll be able to draw on their “understanding” of anatomy to give me a decent guess. LLMs are inconsistent on these kinds of questions and often drop the ball.
We can also point to training data requirements as a discrepancy. In brain organoid experiments, we observe a much lower quantity of examples are required to achieve neural-network-like results. This isn’t surprising to me. Biological neurons have exceedingly complex behavior; it takes hundreds of neural network nodes to recreate the behavior of a single neuron, and they can reorganize their network structure in response to stimuli, form loops, operate in non-discretized time, etc. We don’t know how far transformers will be able to go, but I think that if you want to make a model that holds a candle to that sort of complexity, you’ll need at least many, many orders of magnitude more scale or, more likely, a different, less limiting architecture.
> I mean LLMs crush most humans (even many professionals) at coding/legal/medical/etc. exams.
Exams, yes. Not the work yet. This is also why they've not made doctors and lawyers obsolete: even in the fields where the models perform the best, you're getting the equivalent of someone fresh out of university with no real experience.
I suspect you're right about your point with generalise vs. memorise. Not absolutely sure, but I do suspect so.
I also suspect we'll get transformative AI well before we can train any AI with as few examples as any organic brain needs. Unless we suddenly jump to that with one weird trick we've been overlooking this whole time, which wouldn't hugely surprise me given how many radical improvements we've made just by trying random ideas.
Not quite that bad, but I hear you.
Indeed, but we do also catalogue our cognitive biases — kinda the human version of what are now called hallucinations when LLMs do them. (When Stable Diffusion does it, it's "oh god the fingers").
It's worth caring about both strengths and weaknesses.
IQ tests are averaged out so 100 is average intelligence, not the high point.
The big difference in behavior is in how people or LLMs approach new problems. A LLM is incapable of solving a problem that's similar to one that it was trained on, but with slightly changed requirements. A LLM is also incapable of learning, once you point out their error, or of admitting that it doesn't know.
Regarding people, I find it interesting that even lower IQ people are capable of tasks that are currently completely out of reach for AI. It's not just the obvious, such as self-reflection, but even tasks that should've been solved by AI already, such as driving to the grocery store.
> but people also apply logic and abstract thinking
Which people? If we universally did that as a default, elections would look massively different.
> A LLM is incapable of solving a problem that's similar to one that it was trained on, but with slightly changed requirements
Getting good answers for coding questions on my private code/databases disproves this. The requirements have been changed significantly. I've been through ~20 turn chat with LLM investigating a previously unseen database, suggesting queries to get more information, acting on it to create hypotheses and follow up on them.
> A LLM is also incapable of learning, once you point out their error,
This is the standard coding agent loop - you feed back the error to get a better answer through in-context learning. It works.
> or of admitting that it doesn't know.
Response from gpt: «I apologize, but I'm not able to find any reliable information about a person named "Thrust Energetic Aneksy."»
This looks like it works sometimes, but only if "pointing out the error" is coincidentally the same as "clarifying problem spec". Admittedly for really simple cases those are the same or hard to tell apart. But it always seems clear that adding error correction context is similar to adding additional search terms to get to a better stack-overflow page. This feels very different than updating any kind of logical model for the problem representation.
Updating the logical model of the problem also happens when you do a database investigation I mentioned earlier. There's both information gathering and building on it to ask more specific questions.
You do know that when that happens the LLM usually just throws random stuff at you until you are happy? That is much easier to do than to reason, LLM solved the much easier problem of looking smart than being smart, trick is to make the other person solve the problem for you while attributing it to you.
You see humans do this as well in hiring interviews etc, it is really easy to trick people who want you to succeed.
The model says that because it is trained to say that to specific queries. They have given it a lot of prompts with "Who is X" and showing the responses "I don't know about that person".
The reason you don't see "I don't know" much in other kinds of problems is that there it isn't easy to create such data examples where the model says I don't know while still making the model solve problems that are in its dataset, since it starts to pattern match all sort of math problems to "I don't know" even when it could solve it.
A human can look at his own thoughts and realize how he is solving it, and thus knows a lot better what he knows and not. LLMs aren't like that, they don't really know what they do know. The LLM doesn't know its own weights when it picks a word, it has no clue how certain it is, it just predicts whether a human would have said "I don't know" to the question, not whether itself would know.
They actually do. The information is saved there, you just need to ask explicitly, because the usual response doesn't expose it. (But likely can be fine tuned to do that) https://ar5iv.labs.arxiv.org/html/2308.16175
Also the original claim was that they're not capable of responding with "I don't know", so that's what I was addressing.
All people, otherwise we wouldn't be able to do basic tasks, such as finding edible food or recognise danger.
I dislike how you discount the way other people vote as being somehow irrational, while I'm sure you consider your own political thinking as being rational. People always vote according to their own needs and self-interest, and in terms of politics, things are always nuanced and complicated. The fact that many people vote contrary to your wishes is actually proof that people can think for themselves.
> Getting good answers for coding questions on my private code/databases disproves this.
I use GitHub Copilot and ChatGPT every day. It only answers correctly when there's a clear, and widely documented pattern, and even then, it can hallucinate. Don't get me wrong, it's still useful, but it shows in no way an ability to reason, or the capacity to admit that a solution is out of its reach.
Your experience with coding is kind of irrelevant to the question at hand.
This is simply not true.
Actually, no. We can do 'real' reasoning and come up with novel conclusions. In fact, we have such an example going all the way back to Plato in Meno, about 2300 years ago. It's the doubling of the square dialog of Socrates and the slave boy.
https://classics.mit.edu/Plato/meno.html (Crtl+F for 'square')
Really, the whole dialog is relevant to this discussion with LLMs knowing things.
Source? TFA, i.e. the thing we're commenting on, tried to, and seems to, show the opposite
the total sum of human knowledge is not found in the digital world.