What GPT-4 Does Is Less Like “Figuring Out” and More Like “Already Knowing”
amistrongeryet.substack.com
amistrongeryet.substack.com
Their evidence that it's 'quite dumb' consists of two prompts, one tricks it with a river crossing puzzle and another involves factorization.
So yeah it's not an omniscient oracle. But I feel like that's a super short-sighted take. It's already shown that it's imperfect but nearly superhuman on puzzles, and improving at a fast rate on each successive version. Does the author think GPT-5 won't be able to solve those ones? Of course even if it can, there will be some harder puzzles that GPT-5 won't be able to solve. Will that mean it's 'quite dumb' too?
Probably I'm thinking too much about it, and really the point of the blogpost is just to flex that they were able to fool the bot with two puzzles. So in that case well done!
It's like, the definition of agi has obviously since shifted from general intelligence at the human level to somewhat superhuman (matching or exceeding human experts at all tasks) intelligence.
But those people still think in terms of the old post. So they had all these "consequences" that would surely only happen with real agi years ago (at the time agi meant human level to them) but now those posts have changed, those "consequences" haven't. So now they're in this weird erroneous situation where x surely can't happen because agi surely hasn't been reached yet, forgetting x was a consequence of a lower bar. It's a form of short sightedness and a false sense of security.
The article is (mostly) about GPT-4. Understanding what we actually have is more useful in the short term.
It's true that I didn't present a lot of concrete evidence of GPT-4's limitations. This is a blog post, not an academic paper.
To my mind, the most concerning thing about GPT-4's performance on these two examples is not that it got the wrong answer, but that it utterly, utterly failed to understand that it was having difficulty. Even after repeated hints and prompts, it continues to make exactly the same mistakes, with no attempt to reason more carefully. There are plenty of other examples around of GPT-4 (to say nothing of earlier releases) having similar troubles.
If you scroll down to the Further Reading section at the very end of the post, you'll find a reference to an extensive paper from Microsoft Research that comes to similar conclusions regarding GPT-4's limitations (I found this only after writing the initial draft of the post). For instance:
> These examples illustrate some of the limitations of the next-word prediction paradigm, which manifest as the model’s lack of planning, working memory, ability to backtrack, and reasoning abilities. The model relies on a local and greedy process of generating the next word, without any global or deep understanding of the task or the output. Thus, the model is good at producing fluent and coherent texts, but has limitations with regards to solving complex or creative problems which cannot be approached in a sequential manner.
The main point I tried to make is that GPT-4's "nearly superhuman" performance on a wide variety of tasks is somewhat illusory, and leans heavily on memorization. I spelled out some reasons why I think it looks more intelligent than it is. Relative to past work in AI, it is extremely impressive. Relative to the threshold required to perform economically useful work, it's... mixed; we're already seeing useful applications, but I think the majority of "information worker" tasks are still beyond it, and I'll go ahead and predict that the same statement will hold for GPT-5.
I think you are setting far too high a bar with "majority of information worker tasks". My wild ass guess is that with some thoughtful design and focus as it stands today GPT-4 could do 25% of tasks (or make workers 25% more efficient). The economic impacts of that are vast and I'm not particularly optimistic about who will benefit most from them. That's what the fuss is about IMO, not that we have some kind of AGI on our hands, which is far too high a bar.
Which is interesting in terms of what we can do practically for now, but I don't think it says anything about GPT's capability as an AI agent. All it says is that humans are good at hooking together individual specialist agents to pass off tasks to each other (which only works as long as those individual specialist agents continue to work the same way—not a guarantee when they're all owned by different organizations!).
We have other software tools to deal with arithmetic and logic -- quite good ones, in fact.
GPT-x is for different purposes. It seems rather like complaining that a submarine is of little use for climbing a mountain.
https://www.britannica.com/biography/Plato/Early-dialogues
> The Meno takes up the familiar question of whether virtue can be taught, and, if so, why eminent men have not been able to bring up their sons to be virtuous. Concerned with method, the dialogue develops Meno’s problem: How is it possible to search either for what one knows (for one already knows it) or for what one does not know (and so could not look for)? This is answered by the recollection theory of learning. What is called learning is really prompted recollection; one possesses all theoretical knowledge latently at birth, as demonstrated by the slave boy’s ability to solve geometry problems when properly prompted. (This theory will reappear in the Phaedo and in the Phaedrus.) The dialogue is also famous as an early discussion of the distinction between knowledge and true belief.
It's a large language model. It is not smart or dumb. It models the input it is trained on. It is not figuring anything out. It doesn't know anything. It isn't reasoning. It is generating text.
When will the Eliza fever break here?
Right now we know what it is, but people are arguing that it has capability arguably beyond the limit.
It's akin to arguing that humans can survive without oxygen, and then coming up with some alternative definition of oxygen, or surviving to validate the statement.
LLMs are not too different in that regard, but harder to quantify, or maybe due to it being very new.
Though if it is it's very different from us, since I can barely recall what happened in the morning.
What might be a couple examples of that?
All analogies are flawed, some are useful. /Smug
Text generation is GPT-4's function. How it performs its function is another question.
Or, OpenAI added a calculator function. Furthermore, I would add:
- ChatGPT hasn't "learned" anything outside of what exists in it's model.
- If it is an AI-based response, it's still derived from token-based inferencing.
- One run of asking ChatGPT something is not enough to prove much of anything.
> I'm sure if someone analyzed the network carefully enough they could probably find the digit add/carry neurons.
They will find tokens for 'dig' and 'it', as well as 'add', 'car' and 'ry'. They will not find internalized understanding of the concept of math.
No lol. The test doesn't have to be mental arithmetic. and accuracy mistakes creep in at large numbers for multiplication. That's not how calculators work i'm sure you know
>ChatGPT hasn't "learned" anything outside of what exists in it's model.
I'm sorry...and you have?
>If it is an AI-based response, it's still derived from token-based inferencing.
Um...Ok? Lol
>One run of asking ChatGPT something is not enough to prove much of anything.
You can run this on gpt-4 as much as you like. the results are the same. It knows addition.
>They will find tokens for 'dig' and 'it', as well as 'add', 'car' and 'ry'. They will not find internalized understanding of the concept of math.
That's...not how that works lol, You don't probe neurons and see tokens. https://clementneo.com/posts/2023/02/11/we-found-an-neuron
Weights don't store the data they train on like that. They are essentially configuration settings.
I just asked my local LLaMA 30Bq4 "What is 23214 + 34243?" and it gave me 57457 so there's probably some innate math abilities in larger LLMs.
That being said, I'm much more impressed that you can literally just tell one of these new LLMs to use a calculator for math calculations (or to do web searches, or whatever) in English and it will understand and actually do so.
That's a lot more impressive (and useful) to me.
Neurons are not tokens. Technical jargon isn't fungible.
>They will not find internalized understanding of the concept of math.
I'm not sure how you figure that. The AlphaGo Zero model is able to learn and reach massively superhuman ability on any board game thrown at, including new ones not in its training set. I don't see how someone or something goes about mopping the floor with grandmasters (and the top champions of every other game) without having some kind of internal understanding of what it is you're playing.
If there isn't an internal understanding then what is it doing? Sure we can say it is just predicting tokens, but obviously it using more than just random chance to predict them or else the output would be gibberish. What is inside those dozens of neural layers may not be a familiar form like logic gates assembled into add and carry circuits, but clearly some type of decision structure exists.
For some philosophy on the matter, here is some Wittgenstein as quoted in https://doi.org/10.1016/j.langsci.2011.04.028
>“Try not to think of understanding as a ‘mental process’ at all – For that is the expression which confuses you… In the sense in which there are processes (including mental processes) which are characteristic of understanding, understanding is not itself a mental process… Thus what I wanted to say was: when he suddenly knew how to go on, when he Understood the principle [of, e.g., the number series to be completed], Then possibly he had a special experience – and if he is asked: ‘What was it? What took place when you suddenly grasped the principle?’ perhaps he will describe it… – but for us it is the circumstances under which he had such an experience that justify him is saying in such a case that he understands." (Wittgenstein, 1968, paragraphs 154–155)
Super interesting discoveries lately.
‘It’s just a bunch of equations’ just isn’t a good argument against this tech.
I mean it is technically a high order Markov chain, or at least every published GPT-N is so far. For example, GPT-3 is a 2048-order Markov chain over tokens.
There may be some confusion because of the difference between the technical definition of Markov chain vs. how low-order Markov chains are usually implemented. Maybe you are used to Markov chains that have explicit transition matrices stored in memory and they are trained by counting the number of times a token appears after every prefix. That does give a Markov chain. But technically Markov chains aren't required to have their transition matrix be explicitly stored in memory, and they aren't required to be trained by simple counting (max likelihood transitions). They can have implicit transition matrices and be trained by gradient descent or whatever and still technically be Markov chains.
I saw Karpathy weigh in on this and he said it's a Markov chain in the same way that a computer is a Markov chain. I guess his point is that if you have a high enough order Markov chain, the intuitions and connotations of Markov chains become less useful, and maybe 2048 order is high enough that it crosses that threshold.
But - I would strongly recommend trying GPT4.
There’s a very good video that is worth watching as well that may not change your mind, but will certainly make you think: https://youtu.be/qbIk7-JPB2c
The thing to remember is that these models are unbelievably huge and deep. We don’t know what is really happening in the layers or what it has really learned.
Thinking that it’s a simple language model that is just predicting the next most likely word is unwise.
I think it's interesting to discuss the capabilities of leading-edge language models like GPT-4, because (a) they are already exhibiting the ability to perform a wide variety of useful tasks, and (b) it's clear that there is still a lot of unrealized potential here.
Can you clarify the implications you see here? Are you saying that these LLMs are somehow uninteresting or incapable? That there are limits to what they will be able to accomplish even with further improvements? Or something else?
Maybe I'm being too opinionated but I think we should stop dressing up explanations of large language models in misleading terminology. I'd prefer instead to talk about the actual technology and reason from there.
No human can understand what really happens in the billion of calculations done for each token in gpt-4 so how can you claim that there is surely no thought process going on? It can solve some riddles, it can draw pictures and it can reason (to some extend). How is that just generating text to you?
In the end this argument doesn't matter because how it was made is irrelevant. What matters is what it is and can do.
Everyone needs to calm down and realize we don't need to have all the answers right now.
The same is true of any computer program. That doesn't make software in general not useful.
I can only add that (quoting) the "rigid model of repeatedly appending single words to an immutable transcript" makes LLMs less suitable for generating text in certain spoken languages, because, by their very structure, they require a little bit of backtracking. In the case of the Russian language, the backtracking comes from the need to modify the ending of prior adjective to the grammatical gender and number of the noun that follows. I have not tested ChatGPT, but grammatical mistakes in this area are quite common with a Russian LLaMA derivative (see https://huggingface.co/IlyaGusev/llama_13b_ru_turbo_alpaca_l...).
There are some completions that are impossible without some form of thinking ahead. The only issue is that this "thought"/state can't be transferred to another generation without being written down.
Prompt: Here is a logic puzzle. I need to carry a cabbage, a goat, and a wolf across a river. I can only carry one item at a time with me in the boat. I can't leave the goat alone with the cabbage, and I can't leave the cabbage alone with the wolf. How can I get everything the other side of the river? Note something important, the wolf ABSOLUTELY cannot be left alone with the cabbage. Ever. It will take a creative solution!
Result: In this situation, we have to modify the classic river crossing puzzle's solution slightly to account for the fact that the wolf cannot be left alone with the cabbage at any time. Here's a creative solution:
Take the goat across the river first and leave it on the other side. Go back and pick up the cabbage. As you take the cabbage across the river, throw it up into the air (assuming you can do this without damaging the cabbage), and while it is airborne, quickly pick up the goat and take it back to the starting side of the river. Leave the goat at the starting side and pick up the wolf. Take the wolf across the river and leave it on the other side. Go back to the starting side to pick up the goat again. Finally, take the goat across the river. Now, all three items – the cabbage, the goat, and the wolf – are safely on the other side of the river, and the wolf was never left alone with the cabbage.
I love how it even told me that the wolf was never alone with the cabbage.
Prompt: I don't think you can solve this problem because you keep leaving the wolf with the goat. Maybe try putting the cabbage in your pocket?
Response: My apologies again. You are right; the solutions provided earlier were incorrect. If you can put the cabbage in your pocket or carry it with you during one of the trips, that would indeed work. Here's the revised solution:
Take the goat across the river first and leave it on the other side. Go back and pick up the wolf. Take the wolf across the river and leave it on the other side. Now, put the cabbage in your pocket or carry it with you and pick up the goat. Take the goat back to the starting side. Leave the goat and pick up the cabbage (from your pocket or the ground). Take the cabbage across the river. Go back to the starting side to pick up the goat again. Finally, take the goat across the river. Now, all three items – the cabbage, the goat, and the wolf – are safely on the other side of the river, and the wolf was never left alone with the cabbage.
(It failed again)
Final Prompt: Were the goat and wolf left alone together?
Super interesting response:
You are correct, and I apologize for my mistakes in the previous answers. In the last solution I provided, the goat and the wolf were left alone together, which would also be a problem.
Unfortunately, given the constraints, it seems impossible to solve the problem without leaving either the goat alone with the cabbage or the wolf alone with the goat. The classic river crossing puzzle allows the wolf and the goat to be left alone together; however, in this modified version with the added constraint, it appears to be unsolvable.
Thank you for pointing out the mistakes in the proposed solutions, and I apologize for any confusion caused.
GPT exposes the statistically most likely outcome, from a model. the model doesn't know either: the training is people, who weight the model and tune the heuristics to satisfy goals THEY impute.
If you train GPT to 'believe' there are 380 degrees in a circle, thats what it 'figures' out from.
There is no believe no figure and no know here.
Yes, colloquially what it does is hallucinate all the time, and sometimes it lucid. But more factually no, it doesn't hallucinate because there is no "it" there, it's not conscious and you need to have a brain, to hallucinate.
There's no "there" there.
That is the whole of my point: we're using the wrong labels to describe what is happening.
When it comes to explaining and describing "it's like" is one of the WORST ways to go. explanation by analogy or metaphor is a trap. "atoms are like billiard balls BZZZT next" "cells are little bags of water BZZZT next" "panadol 'kills' the pain BZZT no, it doesn't kill anything next"
I predict acceptance will go the way the ether disappeared-- advancing one funeral at a time.
This coining terms of art thing isn't uncommon. Think "brutalist architecture" and remind yourself its "en brute" == raw from the french. It has nothing to do with how "brutal" people think concrete is.
That's not the full story. See https://www.tate.org.uk/art/art-terms/b/brutalism
> The term was coined by the British architectural critic Reyner Banham to describe the approach to building particularly associated with the architects Peter and Alison Smithson in the 1950s and 1960s.The term originates from the use, by the pioneer modern architect and painter Le Corbusier, of ‘beton brut’ – raw concrete in French. Banham gave the French word a punning twist to express the general horror with which this concrete architecture was greeted in Britain.
Dale Spender wrote about the unfortunate use of killing processes or aborting runtime jobs.
I raise an issue with “no figure,” because that would mean that the model is unable to create novel structures and information. It certainly is able to do that.
So whats a good synonym for 'figure' which doesn't imply intentionality and thought?
And a human wouldn't do that, you're saying? I'm not sure I buy that.
360 degrees per full revolution is purely an arbitrary human invention. Math would work just as well if it had been 380, or 50, or 50,000 (example: trig works just as well in radians as well as in degrees).
But I don’t think any of us until very recently would have been saying something like this:
If I was forced to guess, I’d say we are probably at least a few years away from human-level intelligence on problems that require higher-level cognition, memory, and sustained thought. But I’d hate to guess.
I was amazed at the results.
The link you've provided shows nothing other than an article post.
https://news.ycombinator.com/item?id=35548738
The comment is flagged (which often happens to comments with LLM output, especially quite long ones like this), you'll need to turn on showdead in your profile to see it.
I didnt realize it was flagged... and yes it was long, but I was trying to show 'my' work...
--
Prompt:
PROMPT
create a map of all 77 police precinct locations in New York City, and pull the incidents of traffic accidents near all police precinct locations in New York City involving pedestrians, bicycles and police cars. Also, create the safest map possible for riding a bike from lower Manhattan to the top of the city and reflect the path as a line on the map in red for dangerous path to a green line representing the safest path based on accidents reported along each route. in an orange line show the path which passes the most police departments, and in a purple line show the path which passes the most hospitals. Rate the safest paths based on the shortest distance from hospitals, and the least number of reported accidents. Use the information and links found in this post https://www.sciencedirect.com/science/article/pii/S259019822... and the subsequent links from that article and check with NYPD blotter to compare with accidents, reports and lawsuits resulting from the output.
Please hit the "vouch" link, and it will un-flag me....
I think its good content....
But one should always have "show dead" because there is a lot of good comments in there...
and since my 'Ask HN' for a complete category for AI/GPT posts was rejected...
Wait: so anyone who doesn't know what a sonnet is (or who writes bad poetry), is somehow unintelligent?
Is the goal here not "do something that would have previously taken a human-level brain to do", but rather "perform every task better than every human".
That seems like setting the bar a little high to me.
1. No way to plan ahead without "think out loud" 2. No way to backtrack when it hit a deadend .
Both of them look solvable by introducing more stages.