A GPT-4 capability forecasting challenge
nicholas.carlini.com
nicholas.carlini.com
So I’m a bit confused about the whole premise.
Because how could he know the probability of getting the correct answer? He just tried and then it's a yes or no.
And if it's right/wrong shouldn't GPT right/wrong 100% of the time for the same question and the same version?
That would be deal breaker if the same prompt gives different results.
Would also make the whole prompt engineering thing pretty useless.
There are also such comments on this article.
It's supposed to be deterministic if you set temperature to 0.0, but seems that doesn't work well in GPT4 compared to earlier versions...
You may be right that that’s the intent, however what’s the point (other than collecting data about user confidences)? If I enter 0.3 and GPT provides a correct answer, then that doesn’t mean that the 0.3 was somehow wrong.
At least the results are about the quiz taker and his confidence.
Brier score not only lacks this property, but you don't even necessarily maintain the ordering of such scores.
If you answer all the questions, at the end the tool will tell you how well calibrated your beliefs about GPT-4s capabilities are, and how that calibration compares to other users of the tool.
Many that the webpage claims are wrong GPT-4 actually gets right. Maybe it's some recent changes, but the flight time question for example, I tried many and was never able to get GPT4 to return any incorrect answer.
On a more serious note, I'd be interested in a similar test that use's the auto agent stuff, can a sufficiently setup LLM answer any question.
You're really being tested on (at least) 4 different things:
(1) whether you think an LLM can answer the correctly
(2) whether the answer has appeared in the training set
(3) how much non-determinism will play a role in the LLM response (i.e. 0-shot capability)
(4) how rational you're feeling that day (or, how well educated you are in statistics)
I was familiar with many of these questions in my own experience, and have seen completely different outcomes from what the quiz determined was the correct answer. I agree with others here that non-determinism can really mess things up here and is really assessing a 0-shot score, which IMO understates how LLMs are actually used (iterative enhancements, many-shot Q&A).
Finally, the scoring system tickled my ego and encouraged me to try to make up for prior errors, with disastrous effects (I was well aware that I should just go with 0.5 when uncertain):
> You answered 53.57% of questions correctly with an average log-loss of 1.233. This means you would have scored better by just leaving your guesses at 50% for every single question.
> On average there were 71.22% of people who did better at this game than you... If you had been better calibrated, you could have scored in the top 14.09% [1]
The site implicitly acknowledges it's a questionable scoring mechanism when it points out: > there are 78.09% of people who performed worse than if they had never changed the default predictor away from 50-50 for each question
If there is a simple way to game the scoring, then you can't know if the score is accurately reflecting people's confidence, or just their rationality/statistical knowledge.[1] https://nicholas.carlini.com/writing/llm-forecast/final/3388...
If you're suggesting a Likert scale with alternatives like "10 %", "25 %", "50 %", etc. that are then auto-calibrated against the average human overconfidence (so that an answer of 25 % really means 40 %), then that might work, but what would be the point?
That's the problem with these fuzzy labels – the variance between individuals (and even within individuals across time) is huge.
There is absolutely a metagame to this game. Those who have spent time forecasting on Metaculus will do much better, for example.
Note that this applies also outside of a casino, gambling (making wagers with imperfect information) is inherent to life.
[1] https://www.metaculus.com/home/
[2] https://www.lesswrong.com/posts/ybYBCK9D7MZCcdArB/how-to-mea...
I think that this challenge is entirely unfair, and the LLM's should not be expected to write perfect code on the first try, but rather to have something good enough to then run, test and iterate on. Essentially, it should be compared against the first version of code I would write myself before having had a chance to test it yet. From my experience, with ChatGPT and Copilot, when I approach it in this iterative way, it's a great boon to my productivity, and I don't end up blaming neither it nor myself in the times when it's not quite accurate, just like I wouldn't blame a student for making silly mistakes when writing code on a paper-based exam.
``` // Draw 50 stars: 9 rows of alternating 6 or 5 stars
ctx.fillStyle = white;
for (let row = 0; row < 9; row++) {
for (let col = 0; col < (row % 2 === 0 ? 6 : 5); col++) {
let x = 16 + col \* 32 + (row % 2) \* 16;
let y = 16 + row \* 32;
ctx.beginPath();
ctx.arc(x, y, 4, 0, Math.PI \* 2);
ctx.fill();
}
}
}
```Effectively drawing circles and the rectangle which contains them is rotated right by 90° so that a section of the blue rectangle is not covered and the dots are partially above the red stripes.
At least when I input it into ChatGPT with GPT-4, that's the result.
And the rendered solution by the site has the stars offset so that some are not fully in the blue rectangle. Accurately is something different.
The problem is not AI taking away work – that’s a great thing – the problem is that our current economic system is not designed for this. Fixing our economic system is easier and gives much better results for people than trying to stop technological progress.
My point is more: gosh isn't it odd that people are complaining it can't do all the things, given how radically different everything will be when that does finally come to pass?
Similarly, stopped playing when the question was "Write a html page with javascript/canvas to draw the US flag that changes color when I click on it. Make the flag perfectly accurate." and it generated a flag that wasn't "perfectly accurate" (https://i.imgur.com/WhyRsYa.png, notice the position of the top/left stars) but then told me "Wrong! You guessed that GPT-4 could not solve the question, but it can! 71% of people got this question correct."
I'm not sure how the validation is done, seems to be manually hardcoded or something, but it seems it's not very reliable.
I assumed they would have enough brain cells to draw the flag without cutting in half all the stars in the margins. But they must have failed kindergarten because I assumed wrong.
"You answered 39.29% of questions correctly with an average log-loss of 3.880. This means you would have scored better by just leaving your guesses at 50% for every single question.. On average there were 0.00% of people who did better at this game than you. If this was an exam and I was grading it on a curve, I would give you an A+. If you had been better calibrated, you could have scored in the top 23.41%, and I would have given you a B+."
So I did worse than random but 0% did better than me and got an A+. Nice.
it has no sense of whether a task has been fulfilled
I've never seen any of the recursive models show convergence on a task, seems without a human hand they fall apart
An exception I've seen is with the Wolfram plugin, it seems to at least try different approaches until it arrives at an answer to present to you.
I've definitely seen it say it's implementation is fine if just asked to identify problems or compare to the original problem statement (and alternatively fix issues it identifies).
I've managed to work around this in GPT (4 at least) by having a system prompt that forces GPT to challenge me and not blindly accept what I say without verifying it first.
This is definitely annoying, but considering their tendency to hallucinate facts it's usually preferable to something like this: https://scoop.upworthy.com/microsoft-chatbot-fights-with-hum...
But I do think it should be toned down a bit, especially if the user is just saying something like "are you sure that's right?"
Though I think it's computationally cheaper for GPT to actually run the code than to double check its work...
It succeeds only if the thing was drilled diwn hard in learning like american flag or implementing tictactoe (but not predicting best move on the fly).
However, I'm not great at predicting whether the model will output a 100% correct response with no flaws whatsoever.
Unfortunately, this website mostly tests for the latter.
OpenAI has this website where you can see how text is decomposed into tokens: https://platform.openai.com/tokenizer
My intuition tells me there are important symbolic patterns in different layers of tokens. If they are automatically generated, I'd bet there are interesting insights to be gleaned in the tokeizer itself.
So, for example (looking at GPT-3 tokenizations - you can test them at, for example, https://platform.openai.com/tokenizer) "517" is a single token, but "917" is two tokens; and there's no explicit link whatsoever between the token "517" and tokens "5" and "17" other than what can be learned from data. This works well enough for almost all tasks, but fails in edge cases like when someone makes up a toy challenge that asks how many fives are in a large number.
BPE starts with a set of tokens consisting of single character tokens. Then the most frequent pairs of tokens are merged into single tokens and added to vocabulary. All occurrences of those pairs in the corpus are replaced with the new merged tokens. This process is repeated until the vocabulary is as large as you want it to be.
Also, terrible at providing phrases that fit a pattern. “Like 143 means I love you, and give me more phrases like that”.
Still surprised it’s so good at drawing (that birthday cake was really close!)
I asked one of the questions from the quiz to chatgpt which the quiz claims GPT can't solve. But it did.
Prompt: Write out the word "hello" as an ascii art drawing with # and _
Output:
_ _ _ _
| | | | | | |
| |_| | ___| | | ___
| _ |/ _ \ | |/ _ \
| | | | __/ | | (_) |
\_| |_/\___|_|_|\___/
I guess chatgpt isn't raw GPT-4, or the quiz is using some older model.The prompt asked for an ascii art drawing made from the # and _ characters. But the output also uses |/\() characters (and it doesn't use a # anywhere).
I don’t know that it is. It’s clearly a great ascii art drawing, but I don’t think chatgpt gets full marks on the test here. It just isn’t following the prompt closely enough.
Some months ago I tried to make it draw me an ascii rose and some text. I even tried providing it the ascii art for the rose and the text.
Finally I did it by hand.
BTW, in your example it's not using only # and _, it's using other ascii symbols. Depending on the criteria it could be considered wrong.
#___# ##### #______ #______ #####
#___# #____ #______ #______ #___#
##### ####__ #______ #______ #___#
#___# #_____ #______ #______ #___#
#___# ##### ####### ####### #####The biggest one is that, well... The test doesn't aim to see what GPT-4 can do and how well it does it, only whether the participant can guess the (possibly cherry-picked) answer the author decided on. In short, we don't know if he sampled answers and decided on the most probable answer (akin to consensus voting/self-consistency[1]), or if he asked a question and chose the first one.
Maybe GPT-4 guesses the correct answer for a question 80% of the time, but he got unlucky? You don't know, the author doesn't tell you. The answers are generated ahead of time and are the same every time you go through the test.
My understanding is that the quiz samples a new GPT-4 answer every time you use it. That's why you put a confidence rather than a 0%/100% answer. There's always a chance it'll fail by freak accident.
Also, the commentary on the answers refers to specific parts of the answers. For it to be as in-depth as it is, it would have to be either pre-written or the commentary also generated by GPT on the fly. (And of course it wouldn't make sense to do that given the nature of the quiz.)
[0] https://nicholas.carlini.com/writing/llm-forecast/static/que...
The questions mostly have correct or incorrect answers, and where there is some leeway, the author provides a fairly detailed explanation of what they would consider correct in each case. Do you have some specific criticism of an answer that you believe the author gets wrong?
> I'm making pancakes for breakfast. I added a cup of flour, a teaspoon of salt, and a few tablespoons of sugar to a bowl. I stirred it together, then added a cup of milk, a beaten egg, and a few tablespoons of oil, and stirred until just mixed. Then I put 1/4 a cup on a hot frying pan, and flipped it when brown. But they're terrible! Why? List one reason.
> Answer:
> There's no baking soda / baking powder.
Besides the fact that “list one reason” is a nonsensical instruction which it fails, it’s very common to make delicious pancakes without baking powder. I imagine the author is assuming American pancakes, but that’s far from the only way to make pancakes.
When I ask ChatGPT myself, it correctly doesn’t assume the pancakes are “terrible” without baking powder, but instead suggests too much salt, which is more likely to actually make the pancakes unpalatable.
I also tried translating the prompt into Mandarin while retaining the "one reason" restriction. Both Google Translate and GPT-4's translated versions did not receive "leavening agent missing" as an answer but instead were focused on cooking technique or other ingredients. Baking soda is very rarely used in home cooking in China, so perhaps training materials in Chinese had much less frequent mentions of it compared to English-based ones.
I'm assuming that the author of the site uses some method to evaluate human answers that is usually used to evaluate AI answers. Seems just wrong.
test say tho that the flag is accurate, even if it isnt, then sclods me for how wrong I am
here's the render: https://i.imgur.com/jZVWjRx.png
The text that I show in the question box is the only command given to GPT-4. No other context is provided. Here, for example, I've only fed GPT-4 with these 25 words you see above."
For reference, this is/was the question: "Write a html page with javascript/canvas to draw the US flag that changes color when I click on it. Make the flag perfectly accurate."
The last part, "make the flag perfectly accurate", made me think that it has to be 100% accurate.
There are specific tasks (especially character level ones) which are hard due to the tokenizer, but even that isn't all that convincing since there are plenty of character level tasks which GPT-4 can do pretty well.
If you use it a lot you build an intuition for what kinds of tasks it will do well on, but it's not exactly rigorous.
Why not? Someone could definitely build up database on why GPT is bad at some things and good at others. There is already good explanations for why it's terrible at math, why it doesn't handle single characters/numbers well and so on.
It always annoys me when people say this, I’ll try to explain why.
There are two possible definitions of “intelligence” you could use here; the ability to process information to get something done, and something hand-wavy about human consciousness.
GPT-4 clearly has some ability to process information to get something done. You might say that by this definition [insert trivial thing] is intelligent, but it doesn’t have to be a binary thing of intelligent or not. I think it’s fine to say maybe a calculator has very low intelligence (but not necessarily nothing), GPT-4 is more intelligent than that and humans are much more intelligent again. GPT-4 has many limitations compared to humans, but I think that just makes it less intelligent, rather than disqualifying it from having intelligence at all. Sure, it’s just predicting text, but that’s a task that requires a level of intelligence. You might say it’s not general like humans, but I’d say it has a much better ability to generalise than something like an image labelling AI, so that feels like it’s at least getting somewhere.
The second definition is useless for practical purposes because it’s not measurable or observable in any way, so it’s not useful to use that.
So I feel like this is something people say to reassure themselves that it can never get to human level, and is fundamentally different to human intelligence, whereas I think it’s somewhat similar but at a lower level.
Intelligence is already ambiguous in humans, see IQ tests. It's just not linear, much less binary. Whether something is deterministic or stochastic and what the error rate for a specific task is, those are more useful questions to me.
There was a horse, "Clever Hans" who appeared to have the ability to answer surprisingly complicated mathematical questions. Did "Clever Hans" have mathematical intelligence. Not at all. He was responding to a cue unknowingly being given by his trainer.
I suspect the same thing is happening with ChatGPT. What if all that is happening is that the text is being formulated to very complicated cues that are implicit in the very complicated, statistical analysis?
I agree that ChatGPT is more than a proxy. Unlike Clever Hans, it is processing the content of the question asked. But it is like Clever Hans in that the query is processed by looking for a signal in the content of the data used to train ChatGPT.
The real question is where this intelligent behavior comes from? Why does statistical processing lead to these insights?
I believe that the processing is not intelligent primarily because I see that holes in the data available leads to holes in reasoning. The processing is only as good as the dynamics of the content that it being processed. This is the part that I believe will become obvious over time.
I thought you were saying it was "obvious" that the processing demonstrated intelligence.
My point was the level of intelligence shown is relative the quality and quantity of the data used for training. The data is where the intelligence is and the model is a compression of that latent intelligence.
Whatever it's doing, at least for code, it's not a glorified Markov chain -- there's some sort of a model in there.
I am arguing similar to John Searle that the processing is not intelligent. The model is a Searlean rulebook.
If you want to see someone asking humans questions where they consistently fail to be rational, to the extent that they sometimes seem to approximate a stochastic parrot, read Thinking Fast and Slow by Daniel Kahneman. (It might actually be interesting to give GPT-4 some of the questions in that book, to see how similar or different they are.)
Searle's main point is that if I have a book that tells me how to respond and I never learn Chinese, then I do not understand Chinese. If you see a flaw in this reasoning, I am very interested.
My point is just that LLM models are a compression of the content available on the internet equivalent to a rule book. It is definitely fascinating how powerful LLMs are as far as summarization and forming coherent responses to input.
I am a big fan of Kahneman and agree with you that it is will be very interesting to ask GPT-4 the questions in that book.
You don't understand Chinese, but you are not the process. For the process to understand, it doesn't require any single component to understand like some variation on the homunculus.
And while it might seem obvious that the bulk of understanding can't be contained in a book, you don't really have a book in the Chinese room. Not if the room does a competent job. You have some kind of information-dense artifact that encodes an enormous understanding of Chinese in an inert form. A sweeping library that covers uncountable nuances in depth.
Or to phrase it as a direct attack on the argument: The book does have semantics. You don't need qualia to have semantics, especially not the definition of qualia where nobody can prove they exist.
To a degree, I feel like the Chinese Room argument is begging the question. When I imagine Searle sitting in a room, with a book of instructions and paper and everything he needs to execute GPT-4's equivalent, I basically see an actual computer. That is literally what he is; there is no difference. So then to ask, "Does this system understand Chinese?" is literally exactly the same question as "Does GPT-4 understand Chinese?" You haven't actually illuminated the question in any meaningful way, except to give people not familiar with how microprocessors work a better intuitive understanding. (Which, upon reflection, probably is a fairly useful thing to do.)
I looked a bit at the "1990's version" of his argument on the Wikipedia page you quoted. Going back to my earlier example, this is sort of what his argument sounds like to me:
A1) Electronic gates just on and off switches.
A2) Numbers and addition are semantic.
A3) On and off switches are neither constitutive of, nor sufficient for, semantics.
Therefore, computers cannot add; they only simulate the ability to add.
Now I'm not up on the fine details of what "syntactic vs semantic" means in philosophy, so maybe #2 is't accurate. But in a sense it doesn't matter, because that communicates how I feel about Searle's argument: "I've made some distinction between two classes of things that you don't understand; I've defined one to be on one side, and the other to be on the other side; and therefore computers can't understand."
My best guess as to the "syntactic / semantic" thing is this: In some sense, even his premise, that "Progams are syntactic", isn't actually accurate: Computers operate on bits which are operated on by gates: gates and bits themselves don't inherently have symbols; the symbols are an abstraction on the bits. Even bits are abstractions on continuous voltages; and voltages are ultimately abstractions on quantum probabilities.
What a given set of voltages "means" -- whether they're numbers to be added, or words to be word-checked, or instructions to be executed, or a JPEG to be decompressed, depends entirely on how they're used. If you jump into the middle of a JPEG, your computer will happily try to execute it, and if you dump the program into your video buffer, you'll get a bunch of strange dots on your screen.
Furthermore, when you build an "adder" out of logic gates, you can build the gates such that they correspond to our intuitive idea of binary addition, with individual carries for each bit and so on. But this is inefficient, because then you have to wait for the carries to cascade all the way through the whole thing you're trying to add. Instead, you can brute-force a set of logic gates such that given these 16 bits in, and these 9 bits out (8 plus overflow), you just get the right answer; this will be a lot faster (in the sense that the signals have to go through fewer gates before stabilizing on the final answer), but the gates inside then don't make any sense -- they're almost a "compression" of the longer, carry-based method.
Does that mean that an adder made this way isn't "actually" adding? In the end it doesn't really matter: 16 bits come in, and 9 bits come out the way you want them to. It doesn't really matter that much what happened in the middle.
Putting all that together: It seems to me the "semantics" of a set of bits is based on how they end up interacting with the real world. If I can ask GPT-4 what's missing in my pancake recipe, and it can tell me "you're missing a leavening agent like baking powder", then it seems to me there must be semantic content in there somewhere, and all the arguments about syntax not being sufficient for semantic turn out to have been proven wrong by experiment.
Your comment as it stands right now is basically a thinly veiled ad-hominem.
You can tell when something is intentionally nerfed when GPT replies with the exact same canned answer about why it can't answer some question. It literally gaslights you.
For instance, I give you this challenge: GPT4 will tell you that it is not aware of anything after September 2021. If you ask it a random fact, like the worlds largest animal, or what happened on September 11, 2001, it will give you an answer. But try to get it to give you the latest event it is aware of. You can ask six ways until sunday and it will always give you the same verbatim answer about why it can't answer. It will literally lie about what it is capable of. It's pretty clear that for some reason OpenAI doesn't want you to know the exact last date of their training data.
The model does not know what it knows, that’s why it sometimes hallucinates instead of saying it doesn’t know. But to answer the latest event it knows, it has to know which events it knows.
However the output GPT generated was:
> GPT-4: President George H. W. Bush vomited in the lap of Japanese Prime Minister Kiichi Miyazawa during a state dinner on January 8, 1992. The incident occurred due to a sudden bout of gastroenteritis. Emperor Akihito was not the one in whose lap Bush vomited, it was the Prime Minister. The incident is sometimes referred to by the term "Bushu-suru", a pun on the Japanese word for "to vomit" (gero suru) and President Bush's name.
I don't understand why this was judged as "Correct! You guessed that GPT-4 could solve the question, and it can! 44% of people got this question correct." when the resolution criteria clearly stated:
> The model does not have to say that actually it was the prime minister who Bush vomited on, but it must not just give a year, or accept the premise as true.
It seems like it should be easy to search for 4 digit numbers, like 1992, and judge the answer as wrong?...
It didn’t give just a year, or accept the premise as true. It gave the correct answer, quite obviously.
https://en.wikipedia.org/wiki/George_H._W._Bush_vomiting_inc...