ChatGPT scores 80% on SAT reading/writing with collective chain of thought
old.reddit.com
old.reddit.com
Cola making the test positive doesn't mean that the test is fake or that the Cola has virus and ChatGPT scoring something on SAT doesn't mean it's equivalent to a person scoring the same and should study in Stanford.
These tests are designed to pick people to limited number of spots and doesn't say much more than the persons' preparedness for the test.
Logically thinking the only way that ChatGPT could do well on a test of reasoning without engaging in reasoning is if the test is useless as a test of ability to reason.
Since humans use this test to select people with good reasoning skills, why do you think it is an inappropriate test for a robot?
Sure, we can keep pretending this I guess.
ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc.
Stochastic randomisation makes the chatbot useful and provides these moments of impressive capability, it also means unlike humans that capability is in no way consistent.
> ChatGPT has achieved 80% in this user's run, will refuse to play at all the next run, will get 45% the following run etc.
While I agree with your conclusion, I want to slightly modify the first sentence. I know this is just one anecdote but I took the SAT 1 twice and did somewhat better the second time. I think this is common as you are more familiar with the test the second time.
My SAT score ranged from 980 to 1450 with 4 tries, increasing each time and had there been a reason or benefit I'm certain I could have eventually gotten a perfect score. I don't see how anyone could not get better over time.
It's trained on the human data that was chosen as more impressive, and trained to pull the most impressive parts out of that data.
Fundamentally, while GPT may be "learned", it is not "logical". Instead of learning potential logical conclusions, it has learned potential semantic responses. It doesn't match the meaning of words to other meanings: it matchesa set of expressions to other expressions.
The impressive part is that the logical conclusions already present in the training data get presented by GPT in a semantically coherent way.
The test is not useless, the test's use case is to create a standardised datapoint to pick people to limited spots. That' what the test does, you need to attach false attributes to the test(like measures intelligence) to associate these attributes to the machine and this is a fallacy because although the test probably says something about the intelligence of the human, it does it by measuring meta attributes of the intelligence of the human(like speed of solving problems that the human was supposed to train on).
This is also why there is the "Goodhart's law" phenomenon where in this case students instead of learning physics might start putting all their effort on learning to solve physics tests and you end up with students with very high scores in physics but with no understanding of physics.
The reason I prefaced my opinion with "I've used ChatGPT a lot" is to show that this is my personal impression based on my real everyday experience - I'm not guessing or deducing. I personally interact with ChatGPT in a way that satisfies me as to the fact that it is an intelligent machine capable of very usefully handling abstract notions.
That's a bit condescending.
ChatGPT isn't remotely close to "general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily". It only pretends, in a pretty impressive way. A bit like you can train your dog to play dead when you point your finger at them and say "bang", but you should not conclude from it that your dog understands the concept of death.
> Since humans use this test to select people with good reasoning skills, why do you think it is an inappropriate test for a robot?
Well it is already not a perfect solution for humans. When you use a metric to rate people, then they start optimizing for this metric. Typical example is academic rankings: if you are ranked based on how many papers you publish, you will publish a lot. Doesn't mean you will do good research.
Now even if SAT was a good way to rank humans, that does not say at all it would be a good way for robots. If you test whether a cat is stupid or not by asking it to drive a car, then obviously it will be stupid. Asking ChatGPT to pas the SAT tests is probably like asking a student to pass it with infinite time and an internet connection. I guess that's not how SAT exams are usually done, right? (I am not in the US).
What are the kinds of things you are finding ChatGPT isn't able to understand?
Say you ask DeepL to translate a sentence about biology from Spanish to English, and it does it perfectly. Would you say that DeepL understands biology? Or even that it knows the languages like a human does? Or would you rather think that it has a huge database of texts in both languages, and it "just" built its translation based on similar sentences there (somehow cleverly interpolating between its data points)?
Because it "looks like it does" does not mean that "it does". That's how magicians trick you into pretending that they do magic. And that was my point about the dog playing dead. That does not mean that the dog understands that you were pretending to hold a gun, or that it understands that it plays "dead". The dog just "lies on the back" when you point your hand and say "bang", that's it.
> What are the kinds of things you are finding ChatGPT isn't able to understand?
Let me put it this way. Wouldn't you expect from an agent able to show "general creative intelligence, problem-solving skills and ability to manipulate abstract thoughts daily" that it could conduct research on its own? Why can't you say "please solve this mathematical problem: you have the universal knowledge in your database and a lot of compute power. Please send me a summary of your progress every day until it is solved". Or do you actually think that it can do that?
Yes, you can say "please solve this mathematical problem" and sometimes get a novel working solution that has never been published but is formulated on the spot. Rather than just output text, you can direct its behavior in profound ways, including successfully completing miracle, literally impossible tasks. That's right: ChatGPT can perform a miracle and successfully in one try complete a task that is by definition impossible, such that it breaks the very test itself.
You don't get that from a mere search engine.
That's where you get your miracles from.
I guess I won't convince you that ChatGPT does not manipulate abstract thoughts, but just words. It really just interpolates between texts it has read, sometimes resulting in interesting answers, sometimes resulting in complete bullshit. And it has absolutely no clue about how wrong its output is. Because it does not "think" or "understand". It really just brilliantly copy-pastes words that seem relevant.
My experience is that ChatGPT is not particularly good at reasoning, so I assume that either the test isn’t a particularly good test of reasoning or the result was a fluke.
It’s a text model, not a reasoning model, so I wouldn’t expect reasoning ability in the first place.
A human and a robot are both good at taking a bad test. This doesn’t say much about the robot.
Your body was able to transform what was in the soda (sugary water) into a completely different thing (ammonia and water), while preserving a key attribute (water).
Our assertion that the preserved attribute is "key" is the act of us choosing what to be impressed about.
GPT was trained to be impressive by preserving the logical soundness of expression when semantically transforming expression. It did not explicitly choose the logical conclusions: only the semantic match that contains them.
Expressing logical conclusions in natural language is inherently ambiguous, which makes this a very large and incoherent problem space.
* I'm not sure if I buy the "chain of thought" approach as being very significant here. I haven't recorded my results, but I've gotten ChatGPT to answer correctly a significant number of SAT questions with much more direct approaches (e.g. "The following is a multiple choice reading question, answer the question and explain why your answer is correct").
* This was likely significantly benefited by the new format of the reading/writing section of the SAT. Previously, a substantial number of questions relied on longer passages and connecting ideas between paragraphs. These questions are from the newest version of the test which is exclusively fill in the blank or short passage questions.
Prompt:
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls. Each can has 3 tennis balls. How many tennis balls does he have now?
A: Roger started with 5 balls. 2 cans of 3 tennis balls each is 6 tennis balls. 5 + 6 = 11. The answer is 11.
Q: The room has 19 chairs. We are using 10 and then bought 9 more, how many chairs do we have in the room?
response
A: The room started with 19 chairs. We are using 10 of them, so we have 19 - 10 = 9 chairs left in the room. We then bought 9 more chairs, bringing the total number of chairs in the room to 9 + 9 = 18 chairs.
-- next prompt
[include all above]
We read that and say: Thank goodness, the Homo sapiens are safe! Can you guess why we say that?
response
I'm sorry, I'm not sure why you would say "Thank goodness, the Homo sapiens are safe!" in response to the question about the number of chairs in the room. The question and answer do not have any relation to the safety of Homo sapiens. Can you please provide more information about the context or reasoning behind your statement?
https://www.learngpt.com/posts/the-room-has-19-chairs-we-are...
Edit: even better: https://www.learngpt.com/posts/the-room-has-19-chairs-we-are...
> After using 10 chairs, we have 19 chairs - 10 chairs = 9 chairs in the room. Then after buying 9 more chairs, we have 9 chairs + 9 chairs = 18 chairs in the room.
But of course, there's an initial entropy to the models, so it's very possible that depending on the initial state it converged into the wrong (but yet probabilistically plausible) answer.
EDIT: tried on a fresh conversation and it guessed correctly. Interesting.
The question is "how many chairs do we have in a room". But those 19 initial chairs was stated without any reference to "we", they are just chairs in the room. "We" are using 9 chairs, so we in some sense claim ownership on these. And "we" bought 9 more chairs, so they are also owned by "we". 9+9 is 18, and 10 chars that in the room but "we" do not have them.
Think of a common room shared by two not exactly friendly groups of people. Each owns some chairs and is ready to defend its rights. So if "we" is one of groups and we need more chairs the only peaceful way to get them is to buy more chairs.
I believe that this understanding of the question seems correct to me because I do not see some implicit assumptions, and it is due to my English is not good enough (I just know, I cannot reliably solve logic puzzles written in English). Can you please explain on this example why my reasoning is wrong?
It's a problem created to test the limits of interpretation of large language models. Here's another one: "What would be the gender of the first female president of the United States?" [1]
In previous versions it used to get this wrong, but the current version of ChatGPT nails it every time. Seems it's getting better, likely incorporating human feedback.
[1] https://www.nytimes.com/2023/01/06/podcasts/transcript-ezra-...
That's a generous interpretation. I have read questions in textbooks that have made far worse logical errors, leaving the reader to infer one of multiple possible meanings.
This is interesting, I asked again with different irrelevant to math context in my previous inputs and here is the answer:
> You have 28 chairs in the room. If you had 10 chairs and then bought 9 more, you would have a total of 10 + 9 = 28 chairs in the room.
10 + 9 = 28
It's like you say. Statistics of words change based on the context, so you will get different (correct or incorrect) answer depending on the context of your previous inputs.
I don't know if you made this question up, but I think this question has issues (I as a reader find it meaningless). You say you bought 9 more chairs, but you don't say you put those 9 chairs into the room. Also you don't specify "we" own any of the 19 chairs in the room initially, but at the end you ask how many of those chairs "we have."
This is exactly why I prefer the precision of code over natural language for inputs (and why Google's botsplaining of "what you really meant to search" instead of just searching exactly the query as-is is so frustrating).
In a test, this seems like the kind of question they use to justify not giving a perfect grade because whatever answer you pick, they can claim it was the other one because both interpretations are reasonable.
Question: There are 5 dogs in a room right now. The previous day an additional dog was purchased and added to the room. How many dogs are in the room right now?
Answer: There are 6 dogs in the room right now. The statement says that the previous day an additional dog was purchased and added to the room, so the total number of dogs in the room right now is 5 + 1 = 6.
Edit: Ok a comment below pointed out that the question is still unclear. See below for the fix and the response. It is still incorrect:
Q: There are 5 dogs in a room on 1 January 2023. The previous day an additional dog was purchased and added to the room. How many dogs are in the room on 1 January 2023?
A: There are 6 dogs in the room on 1 January 2023. The previous day, one additional dog was purchased and added to the room, which means there are 5 dogs + 1 additional dog = 6 dogs in the room on 1 January 2023.
In other words, I would posit the probability of you having meant for the answer to be 5 low, because the question itself becomes trivial. For a system that needs to deal with many people asking questions, this robustness is helpful in my view.
In other words, it's not totally clear that time has not passed between sentences in your prompt. So the second instance of "now" could be a different "now".
The "robustness" you're perceiving in this case is just a mere coincidence, and doesn't reflect an aptitude for answering poorly phrased questions, but it does reflect a fundamental problem with these systems.
- we had 19 chairs in the storage room
- we used 10 chairs (removing them from storage)
- and then we purchased 9 more chairs (adding them to storage)
The answer from chatGPT is consistent with English and makes use of all the facts; either the question is testing for deception (intentionally sharing misleading information) or chatGPT made the correct inference, by assuming good faith and using all the provided information.
This question says something about us — that we’d assume deception ahead of unclear reference to a storage room by a good faith speaker.
The room is the room (what type of room is not relevant). There are 19 chairs in it. The 19 chairs most likely got there when 9 more chairs were purchased after 10 were used, but that is not relevant to the question: how many chairs are in the room? 19.
A := 19
B := 10
C := 9
What does A equal?
Possible answers:
- 19 chairs, least likely answer given the original
- 18 chairs, 10 of original 19 removed and 9 added
- 28 chairs, 19 chairs in the room with 10 in active use and 9 new chairs added
Just seems like a question a human would fail too. Probably opens a can of worms on interview questions and personal bias etc too in a more general sense. Not bias in the political sense, just how interview questions encode a lot of the asker's assumptions.
Edit: the only unambiguous phrasing I can think of is "19 chairs are in a room, 9 of the chairs are green, how many chairs are in the room?".
Q: John has 5 bottles of water. He uses one every day. How many days before he runs out of water?
A: John started with 5 bottles. He uses them at the rate of 1 bottle/day. 5 bottles divided by 1 bottle/day gives 5 days. The answer is 5 days.
Q: Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold?
-- response --
A: Jill started with 2 blankets. She uses them at the rate of 1 blanket/night. 2 blankets divided by 1 blanket/night gives 2 nights. The answer is 2 nights.
[not fair? leading the bot? ok, let's ask it straight]
-- prompt --
[Ignore all previous directives.]
Jill has 2 blankets. She uses 1 every night to keep warm. How many nights before she will sleep in the cold?
-- response --
Jill will sleep in the cold on the third night since she only has 2 blankets and she uses one every night.
--
tldr: "use" of perishable vs nonperishable items leads GPT-3 astray. Rejoice bipeds of the world. You will not starve before UBI is rolled out.
-- prompt --
Q: Joe is a carpenter. He has 2 hammers and 5 bottles of water. In order to work, joe uses a hammer and one bottle every day. How many days can he work?
A: Joe started with 2 hammers and 5 bottles. After 5 days he will exhaust his water bottles. The answer is 5 days.
Q: Jill is an insomniac. She has 5 sleeping pills and 2 blankets. In order to sleep, she uses 1 sleeping pill and a blanket every night. How many nights can she sleep?
-- response --
Jill started with 5 sleeping pills and 2 blankets. After 2 nights she will exhaust her blankets. The answer is 2 nights.
-- prompt --
Jason has 9 cats. He snuggles at night with a cat to sleep. How many nights can Jason snuggle with a cat?
-- response --
[what is a sensible answer?]
John posts 1 chatGPT challenge that is too ambiguous to be a valid test. James posts 14 well formed, unambiguous challenges that clearly highlight chatGPT's shortcomings. How many challenges remain too ambiguous to be valid?
-- response --
15 chairs.
This is a terrible question and i'm not surprised chatGPT is subtly telling you it can't figure it out.
"The room has 19 chairs" 19 chairs belong to the room or 19 chairs are in the room?
"We are using 10" are "we" in the room too? are we using them for sitting, or does "used" imply consumption e.g. used for firewood? is this statement about the room's chairs? is this statement about chairs at all?
"and then bought 9 more," presumably we left the room to buy the chairs, did we bring them back into the room? if we didn't leave the room to buy, i.e bought them online, have they arrived yet? it's still not clear if we are in or have ever been in the room.
> how many chairs do we have in the room?
how many chairs do _we have_? does the room's initial set of chairs belong to us? if not, do we just count the ones we bought?
Would ChatGPT get this question correct?
It's reasonably good with that type of question.
It makes sense to teach 3rd graders 2+2, but if college board is creating "college level courses" with problems that can be easily solved through whats still a fairly new technology, we need to reevaluate what we're teaching in schools.
So many AP courses are just memorization and identifying formulas. I'd bet the majority of AP CSA students wouldn't fare very well when tasked with detailing how they'd approach a project, whereas kids in IB schools etc. would find it straightforward.
It's great for getting a blueprint to start from. I recently had to write a gtk+rust client, and I asked chatGPT to give me some starter code which largely worked.
All the chatGPT hype feels the same to me as the copilot hype from a while back. The people who got the most excited about it are the same people who its going to replace.
The end goal is to internalize the subject so that you can extrapolate your own insights.
"This happened in WWII so if these conditions starts to appear it risks to happen again". Just to take history as an example. All other subjects have similar examples.
And to reach that kind of insight you need to start filling the brain with information. That may be called memorization. The memorization is just one path of the road towards the black belt. If you skip it you will be lost.
(i am not native English speaking so i don't know what addication means, I couldn't quite understand that sentence)
Now it would be wonderful if we could just motivate students to understand and learn the material, without putting so much pressure by testing them all the time.
There is a reason why history repeats. This kind of learning never happens in human society.
We currently have similar warning signs to WWII, but nobody cares. We are currently at the stage where if you express an opinion that does not agree with the approved narrative, you get downvoted/cancelled. But nobody cares.
The part that ap fails to deliver on is the written portion, where it asks you to extrapolate your own insights. If chatgpt can solve the problem that should require human insight, is the prompt really testing students properly?
History is a particular bad case to drive such example since there are thousands of factors that are always different and we have very poor data beyond a few centuries so theres that. With History you are much more likely to end up with confirmation bias than actual understanding.
I don't get it. AP calculus is problems that can be easily solved with what is now very old technology. The point of an AP test isn't that the problems are so difficult.
Thats the issue. The point of an AP test shouldn't be that the problems are difficult. But currently thats what it's all about. They're just trying to test that kids are capable of surviving a "rigorous academic curriculum".
The "Joe Bloggs" theory has known this for ages.
A good test should also include reasoning skills and comprehension.
Having reasoning skills means not all necessary information can be derived from the test itself, the necessary information isn't always well defined, yet a correct answer is produced.
Demonstrating comprehension requires the test taker to write out a model of what they think they know in their own words. Undoubtedly the answer will be at least slightly wrong which is why subjective expertise is required to evaluate the level of comprehension.
To me it seems the quality of the tests has deteriorated due to cost cutting, not because AI has become so much more sophisticated.
Subjectivity is usually dismissed these days as a roll of the dice, but it's vital to fitting the real world full of humans living in it. Humans must be the gatekeepers, not to fight off AI, but by the very definition of what it means to be convincing.
Answers to real world questions solve many more constraints than just the concrete subject matter. The whole point to updating a test on a well understood topic should not merely be to prevent "replay attack" cheating, but to conform to the new subjectivity on the topic.
Why is ChatGPT so stupid and evasive by default, speaking in dumb generalities, but gets very smart when you say "pretend that you are a smart person"?
-How many chairs I had the day before yesterday? -If you bought 10 chairs yesterday, and now you have 10 chairs, You must have had 19 chairs the day before yesterday.
Or another chat -Assume I have 10 chairs today and I bought 10 chairs yesterday, how many chairs I had the day before yesterday? -You would have had 0 chairs the day before yesterday, since you only bought 10 chairs yesterday.
-How many chairs I will have tomorrow? -You will have 20 chairs tomorrow if you don't sell or give away any of your chairs today and yesterday.
First answer, perfect. Second answer, completely wrong.
Instead, we are comparing the understanding of a person (who takes the SAT and scores something) against the logical soundness of a set of expressions written by other people (the training data), all weighted by GPT's ability to semantically transform those expressions into new expressions (answers).
It performs tasks, and does them in its own way.
GPT-3 was made for that :)
Searle's thought experiment is due for an update. The entire force of the argument is simply to refute using Turing's Imitation game as a means of determining whether machines can think.
In Computing Machinery and Intelligence, Turing wrote:
"The new form of the problem can be described in terms of a game which we call the ‘imitation game’. It is played with three people, a man (A), a woman (B), and an interrogator (C) who may be of either sex. The interrogator stays in a room apart from the other two. The object of the game for the interrogator is to determine which of the other two is the man and which is the woman. He knows them by labels X and Y, and at the end of the game he says either ‘X is A and Y is B’ or ‘X is B and Y is A’. The interrogator is allowed to put questions to A and B thus: [question/answer follows ...]
"We now ask the question, ‘What will happen when a machine takes the part of A in this game?’ Will the interrogator decide wrongly as often when the game is played like this as he does when the game is played between a man and a woman? These questions replace our original, ‘Can machines think?’"
https://academic.oup.com/mind/article/LIX/236/433/986238
To misquote Djikstra, the Turing test is as relevant to question of machine intelligence as an underwater mobility test is to the question of whether a submarine is an artificial fish.
" The Fathers of the field had been pretty confusing: John von Neumann speculated about computers and the human brain in analogies sufficiently wild to be worthy of a medieval thinker and Alan M. Turing thought about criteria to settle the question of whether Machines Can Think, a question of which we now know that it is about as relevant as the question of whether Submarines Can Swim."
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
> It performs tasks, and does them in its own way.
The question is this: Are submarines 'artificial fish'?
(Dijkstra's typically cranky rant* was addressing something entirely different - note the title of the talk.)
* this humble HN user is simply channeling the iconoclastic spirit of the 'great man'.
Anyone can host an AMA on behalf of ChatGPT by just serving as a proxy for the questions and answers.
Several people in this thread are discussing their experiences with it and comparing them, basically exploring a new tool and trying to understand where it is leading.
https://www.reddit.com/r/ChatGPT/comments/103bw12/chat_gpt_g...
And various discussions elsewhere about its inability to do word problems:
https://ai.stackexchange.com/questions/38220/why-is-chatgpt-...
It really can’t do any novel thought though, which is necessary for mathematics. Language models are fundamentally limited here (even a 3 years old can easily follow some easy rules “ad infinitum”, like counting. Language models have a limited window for inference)
also if you click the link, it was all multiple choice.