It's cool and it will probably be able to get this right one day but it's a big goal to miss.
It's cool and it will probably be able to get this right one day but it's a big goal to miss.
If wrong answers still count because humans sometimes make mistakes, then I guess it won’t be too difficult to construct an impressive mathematical AI.
It's very tempting to give these systems the benefit of the doubt, but that tends to lead to hugely inflated conclusions about their capabilities. Remember that something as simple as ELIZA was perfectly capable of fooling humans who were predisposed to believe it was intelligent.
The system fundamentally cannot do this. You can make it generate text that is like what someone would say when asked to show their working, but that's a different thing.
> It's very tempting to give these systems the benefit of the doubt, but that tends to lead to hugely inflated conclusions about their capabilities.
I agree. I am seeing a bit too much over-optimistic predictions about these things. And many of these predictions are stated as fact.
Yes, we've had Wolfram Alpha for ages. For me, the biggest problem Wolfram Alpha is that it often doesn't understand the questions, and while I also sometimes get that with ChatGPT, the latter is much much better.
I've not had a chance to play with the plug-in that connects GPT to Wolfram Alpha.
> It's very tempting to give these systems the benefit of the doubt, but that tends to lead to hugely inflated conclusions about their capabilities.
This is an excellent and important point.
I think people who treat it as already being superhuman in the depth (not merely breadth) of each skill are nearly as wrong as those who treat it as merely a souped-up autocomplete.
I've only played with 3 and 3.5 so far, not 4, my impression is that it's somewhere between "a noob" and "genius with Alzheimer's".
A noob at everything at the same time, which is weird because asking a human in English to write JavaScript with all the comments in French and following that up with a request written in German for a description of Cartesian dualism written in Chinese is not something that any human would be expected to do well at, but it can do moderately well at.
Edit: I should probably explicitly add that by my usage of "noob", most people aren't even that good at most things. I might be able to say 你好 if you are willing to overlook a pronunciation so poor it could easily be 尼好 or 拟好 (both of which I only found out about by asking ChatGPT), so I'm sub-noob at Chinese.
I do think that this specific moment we're going through, in early 2023, is producing some of the most fascinating, confidently incorrect misunderstandings of chatgpt, and I hope that someone is going through these comment sections and collecting them so that we can remark on them 5 to 10 years down the line. I suspect that in time these misunderstandings are going to be comparable to how boomers misunderstand computers and the internet.
I recently asked ChatGPT to write me some Go code. What it wrote was mostly fine, but it incorrectly used a method invocation on an object, instead of the correct code which was to pass the object to a function at the package level (aka static method).
I think it's a real stretch to suggest that this could happen as a result of simply regurgitating strings it's seen before, because the string it emitted was not ever going to work. To me, it looked for all the world like the same kind of conceptual error that I have made in Go, and the only way I could see this working is if GPT had a (slightly faulty) model of the Go language, or of the particular (and lesser known) package I was asking about.
It's felt more like it "forgot" - like me - how the package worked, so it followed its model instead. That error was WAY more interesting than the code.
And this is the error I notice people very consistently making when they evaluate the intelligence of ChatGPT and similar models. They marvel at its ability to produce impressive truths, but think nothing of its complete inability to distinguish these truths from similar-sounding falsehoods.
This is another form of the rose-colored glasses, the confirmation bias we are seeing at the peak of the current hype cycle, reminiscent of when Blake Lemoine convinced himself that LaMDA was sentient. A decade ago, techies were dazzling executives with ML models that detected fraud or whatever as if by magic. But then, when the dazzling demos were tempered by the brutal objectivity and rigor of the precision-recall curve, a lot of these models didn't end up getting used in practice. Something similar will happen with ChatGPT. People will eventually have to admit what it is failing to do, and only then will we start up Gartner's fabled Slope of Enlightenment to the Plateau of Productivity.
RLHF and the tokenizer together explain many of the more common failure modes.
The way to test that, of course, is to give it problems that it hasn't seen. Unfortunately, because GPT has seen so much, giving it problems it definitely hasn't seen is itself now a hard problem. What's been shown so far: OpenAI's benchmarking is not always rigorous or appropriate, and GPT's performance is brittle and highly sensitive to the problem phrasing [1].
I agree with the article that GPT's training enables it to access meaningful conceptual abstractions. But there is clearly quite a lot that's still missing. For now, people are too excited to care. But when they try to deliver on their promises, the gaps will still be there. Hopefully at that point we will embark towards a new level of understanding.
[1] https://aisnakeoil.substack.com/p/gpt-4-and-professional-ben...
If a human mathematician said the things ChatGPT said in this dialogue, you would wonder if the person had recently suffered a severe traumatic brain injury.
There is a causality inversion here.
The only reason people know it can't do this is because they have tried and seen it cannot do this.
We do not have very precise bounds a-priori what GPT can and cannot do. We only learn them from black box testing.
No mathematician would make this exact mistake because it looks ridiculous, but for example I made similar mistakes multiple times, claiming low numbers prime or not prime wrongly during a reasoning. Especially when working at night.
That in no way precludes deriving deep understanding of some other concepts.
Even humans are not expected to be able to learn anything we throw at them.
https://ar5iv.labs.arxiv.org/html/2201.02177
This paper makes me believe it's less about the lack of ability to understand maths, but the power to learn more of it.
I don't think it really makes the case that you think it makes. Humans learn language and concepts through sentences, and in most cases semantic understanding can be built up just fine this way. It doesn't work quite the same way for math. When I look at the numbers in the example, I have no idea if they are prime or factors because they themselves don't have much semantic content. In order to understand whether they are those things or not actually requires to stop and perform some specific analysis on them learned through internalizing sets of rules that were acquired through a specialized learning process. Humans themselves don't learn math by just talking to one another about it, rather they actually have to do it in order to internalize it.
I suspect it's simply a case of we haven't sent ChatGPT to highschool yet. There will be some particular trick in terms of training methodology that trains the network to perform that more specialized analysis that doesn't simply emerge out of 'predict the next token from all this random internet text' as it's likely severely underrepresented in the training data in the first place, but I think it ought to be perfectly possible.
98.5% accuracy on addition arithmetic this way with GPT-3
Or did you use a calculator to verify?
Because that would be exactly what ChatGPT with plugins does.
As always, anthropomorphize an LLM at your peril. These things are not people. No one understands how they work.
The first thing I did was ask it about the prime factorization of 3,442,177,452. It said it was 2^2 * 863 * 198853. Then I asked it to compute 4 * 863 * 198853. It said:
> the product of 2^2 * 863 * 198853 is equal to 3,442,177,452, while the product of 4 * 863 * 198853 is equal to 689,401,672.
I then got it to note that 2^2 = 4, and questioned it about what this implies about its previous statements. After a couple attempts, it did figure out the results of the two multiplications should be the same, and that therefore there was a mistake in its previous responses. It then claimed that the results of both multiplications should be 3,442,177,452. Then I asked it how it derived the answer. It said:
> I performed the multiplication of 4 * 863 * 198853 using the multiplication algorithm that is part of my programming. This algorithm uses basic arithmetic operations to calculate the product of the three numbers, resulting in the correct answer of 3,442,177,452.
Then I told it that I got 4 * 863 * 198853 = 686,440,556 from Google's calculator. At this point, it disclaimed all previous results. I then asked it why it had gotten incorrect answers, and got this astonishing response:
> The reason for the error is that I incorrectly calculated the product of 4 * 863 * 198853 by mistakenly assuming that 2^2 is equal to 2, rather than 4.
> In my first response, I correctly calculated the product of 2^2 * 863 * 198853 to be 3,442,177,452. However, in my subsequent responses, I made a mistake by stating that the product of 4 * 863 * 198853 is equal to 3,442,177,452, which is incorrect. Instead, the correct product of 4 * 863 * 198853 is 686,440,556.
Recall that I had previously gotten it to agree that 2^2 * 863 * 198853 and 4 * 863 * 198853 were equal because 2^2 = 4. Here it immediately goes back on that in a slapdash attempt to explain its previous mistakes.
What I find interesting about this whole AI discourse is that no one seems to bother with this kind of test, one where we take the hype about GPT's "general intelligence" seriously and actually test the limits of that, with true respect for the concept of "general intelligence". Honestly, there seems to be a sort of soft bigotry of low expectations. Everyone wants to give the AI problems that it can solve and pat it on the head for its successes. There is less interest in finding ways to probe whether it is truly as "generally intelligent" as the hype says.
In other words, mathematics or arithmetic is not highly encoded in language. It's not that nobody can think of these tests, it's that they don't say what you imagine they do. A poor understanding of math is simply that...a poor understanding of math. General understanding is not binary. You can understand some things well and not understand others.
That is one. 2, people really need to start doing these gotcha tests on GPT-4. It's just much better across the board. And has a much better understanding of arithmetic than chatGPT.
ChatGPT is often good at understanding patterns involving the substitution of one string for another. So you might hope that it could do well in a case like this. But it doesn't really. It is aware of the laws of arithmetic and can explain them in the abstract but it can't apply them consistently in the real world.
I look forward to seeing how GPT-4 does as well. I don't have ready access to it. Looks like I would have to pay to get ChatGPT 4. But I will go out on a limb and predict that it won't be hard to generate this kind of issue with the new version.
The primes thing was just an illustrative example.
You want it to do large scale arithmetic with high accuracy? Describe arithmetic as an algorithm to be performed on two numbers. https://arxiv.org/abs/2211.09066
It's not about chatGPT not understanding anything. It's about chatGPT not understanding math very well. That bleeds into understanding of related concepts as well. LLMs build all these models of the world from the text they train on. Well not all if it is accurate.
If GPT were generally intelligent, we shouldn't need to devote a special research project to teaching it math. We could just throw a math textbook at it, explanations and worked examples, and it would figure it out from there. Almost certainly its training data contains a great deal of such material already. That this doesn't work suggests its mental architecture is insufficient to grasp what it's been told. (Note that people are not advantaged with any specialized symbolic representation of numbers, like the integer data type a computer has. We manipulate numerical symbols as text, same as the AI.)
It's all well and good that it can improve when a special effort is made, but it sounds like even with that special effort, it still doesn't show the level of competence one would expect from a human-level intelligence with access to virtually infinite, untiring silicon computational resources.
GPT has access to abundant materials to learn the laws of arithmetic from, and it can tell you what they are (because it memorizes everything) but it isn't really understanding what it's learned. That points to a shortcoming of the architecture that won't be solved by merely throwing more data at it.
Do people not explain things they don't fully understand?
Understanding is not binary.
This is kind of problem I keep seeing. Expectations and post shifting have grown so much that a significant chunk of the human population wouldn't even pass so called General Intelligence requirements.
There's no post shifting. The research community has been setting itself realistically attainable benchmarks. Now that the research community has made a lot of progress against its benchmarks, we have hype, claims of general intelligence. Which attracts people like me, who compare the hype to actual performance. And as I said elsewhere, the performance of GPT on the questions I posed is only comparable to a human with a severe traumatic brain injury.
Dunno what to tell you other than textbook and practice problems is far from the solution you think it is for a big chunk of the population.
Why would training a neural network architecture that lives and breathes symbols somehow yield an entity with intelligence like that of an average human? Most humans learn primarily from completely different sources.