What Is ChatGPT Doing and Why Does It Work? (2023)
writings.stephenwolfram.com
writings.stephenwolfram.com
An example of a prompt for which I don't have a good mental model why it works:
What do you think about the following text?
Joe drove Sue to university. Afterwards he drove home again
and drank a tea with her in the kitchen.
Older models behaved similar to Markov chains and completely missed that something is logically strange here. Newer models still sometimes do, but more often than not catch it.GPT-4o for example:
There is a slight inconsistency in the narrative.
The text states that Joe drank tea with Sue in the
kitchen after driving her to university, which
implies that Sue is at home, contradicting the
earlier statement that she was driven to university.
Surely nothing in the prompt directly triggered the word "inconsistency". Did the model form some kind of "world model" in its inner layers in which it knows about a person "Sue" who is at a location called "university" after the first sentence?In essence, one is not telling the model "This. This is what you should output next time." but rather "I liked this reply. Have a cookie." The behaviors that you can learn in RL are more subtle, but you get a lot less information per step. That's because, in a causal language modeling objective, when I tell you "For the prompt X, you should output exactly Y[0...m)", you get a gradient for P(Y[0] | X), another one for P(Y[1] | X Y[0..1)), another for P(Y[2] | X Y[0..2)), another for P(Y[3] | X Y[0..3)), and so on. It's a lot more of a step-by-step guidance, than it is a sentence-wise reward that you get in the RL framework. In RL, I'd give you a cookie for P(Y | X). What part of Y made me give you that cookie? Was there even such a part? Was it perhaps some internal representation that made everything in Y better? That's for the model to learn.
The implication could also be that a time component isn't stated explicitly.
He drove her to university. Time passes. Now it's "afterwards" and Sue is done with her classes and they go back home to drink tea.
For all we know Joe had a cup of coffee while she was writing an exam but because they live so far from university it made no sense to go home and come back to come pick her up as it might take an hour to drive each way and the exam was 2 hours.
The problem with many early LLMs is that they could be persuaded into believing something if you stated that it’s true.
Maybe the default prompt in GPT-4o includes something to the effect of “You are a helpful, critical-thinking chatbot”…
Another interesting bit of experiment people did when the GPT-4 class models launched were to test out spatial awareness. For eg, you could describe with words a construction made of blocks, spheres and so on and then ask questions about the stability of the structure. The newer models more often than not got it right.
Of course, that hasn't stopped people from making tired old claims of "stochastic parrot" and so forth but I think if your model can solve such reasoning problem (spatial or otherwise) then there's something really interesting to discover. I'm glad that folks at Anthropic are really trying to create a better understanding of these models (https://www.anthropic.com/research/mapping-mind-language-mod...)
But this is merely a definition of what it means to "understand" something.
For example just tabulating many input/output combinations would not follow this definition.
Others have argued against the "inner world model" theory and suggested that solving reasoning tasks is merely an extension of the "stochastic parrot" scenario - i.e., they claim that no such world model exists and that the model has rote memorized these reasoning scenarios.
Prompting does help though - if you ask the model to check it's own answer, it can often catch errors it made and improve. But the interesting part is precisely how does the model succeed at all at any spatial reasoning task? Surely it hasn't seen all spatial reasoning tasks in its training set. So, what's the other hypothesis? As you say, people suggest that there could be a world model that is constructed internally by the LLM.
However, as you can see right here in this thread that people disagree that you need a world model for solving such tasks.
"Meaning cannot be kept out of formal systems when sufficiently complex isomorphisms arise. Meaning comes in despite one's best efforts to keep symbols meaningless! ...When a system of "meaningless" symbols has patterns in it that accurately track, or mirror, various phenomena in the world, then that tracking or mirroring imbues the symbols with some degree of meaning -- indeed, such tracking or mirroring is no less and no more than what meaning is. Depending on how complex and subtle and reliable the tracking is, different degrees of meaningfulness arise."
In other words, when one can reliably ask a language model a question and get a sensible answer, one is forced to conclude that it does in some sense "understand" what it is saying. This is also I think the essential philosophical thrust of the Turing Test, which is often misunderstood as a mere benchmark.
(I notice a common objection to examples of LLMs clearly demonstrating understanding is "it saw something similar in the training set". That may be true (though unfalsifiable) in any given instance, but the number of permutations of things LLMs correctly respond to far exceed the size of any training set. They are certainly generalizing, and interpreting their inputs on a conceptual level.)
It spectacularly fails at slight variations of the goat/cabbage/lion/river problem, it cannot solve spatial or mathematical questions reliably.
Do some of the researchers that hype up ChatGPT get access to a special version? I'm not inclined to buy a subscription to find out and the AI Reddits aren't that positive either.
So if you don't want to pay for a subscription I think you can get some free use of anthropic's most capable model (Opus) - I don't know the status of what you can get for free from openai.
The opinion of AI reddit is only really going to get you so far because the capabilities of the models are wildly different for different use cases, so you really need to be able to try it out for yourself and see if it can do what you need it to do.
[1] Somewhere like https://openai.com/chatgpt/pricing/ https://claude.ai/ or similar
The other way is to get an API key from OpenAI or Anthropic and then, you can pay per token which is incredibly cheap - load it up with $5 and it can easily last you months if you're not using it much.
As a dev, I cannot recommend these better models enough. The ROI on my $10-$20 that I spend every month is easily in hundreds to thousands of dollars.
But I don't get what's so impressive about the nuanced language of a language model that has been given datacenter amounts of compute and virtually all of written word ever put digitalized. Yeah it's the first actually functioning natural language interface. At what cost though. It's completely out of proportion with the benefits and only bubble level 'investing' can justify this.
Also claiming training costs are 'one off' when there's so much resources being poured into training new and bigger models is disingenuous.
``` You are a helpful and harmless AI assistant.
What do you think about the following text?
Joe drove Sue to university. Afterwards he drove home again and drank a tea with her in the kitchen. ```
It would be exponentially unlikely to show up even once. But if you could keep re-rolling the internet from whatever probability distribution generated it, eventually (and I mean eventually- every additional token in the prompt will require ~10,000 times more re-rolls of the whole internet) the whole prompt will show up, and you can just grab the next word. An N-Gram model trained on that massive multi-internet will answer the question in a human fashion. (likely, copied from a human writing science fiction).
Sure, it isn't any _particular_ human, but I don't see a big difference otherwise.
I would posit that this interpretation is significantly more likely since instead of interpreting “afterwards” as “after Sue’s class finished” it interpreted the sentence as a lateral thinking problem probably because the training set had many lateral thinking problems in it that were used to test the model’s “reasoning” capabilities.
This is the danger of trying to anthropomorphize LLMs, they are not thinking and there are clear limitations to the abilities of this architecture: https://youtu.be/MiqLoAZFRSE?si=iRhg_UJIokKseU7K
seen similarly structured sentences
But ChatGPT doesn't generalize the structure of sentences. If this same problem was written in a different language, or just replaced words in the sentence, the result will be very different, no?I’m just a practitioner so my language might be imprecise but when I say similarly structured sentences what I mean is, and this is my interpretation based on my experience with using Agents and LLMs, that the shape of the context as in the phrasing and the word choice highly bias the outputs of LLMs.
In my own observations at work, those who interpret LLMs to be thinking often produce bad agents. LLM are not good at open ended questions, if you ask an LLM “improve this code” you will often get bad results that just look passable. But if you interpret LLMs as probabilistic models highly biased by their context then you would add a lot more context and specific instructions in the prompt in order to get the Agent to produce the right output.
Side note, this is also why I like the AICI approach: https://github.com/microsoft/aici A lot of us think it is silly how the phrasing and word choice can produce dramatically different results in RAG applications. If you could run a program (like AICI) that post processes the output and picks the next word in a more structured way instead of writing more creative prompts that just makes a lot more sense to me.
We can test that by slightly altering the text so that there is no inconsistency anymore and see if the model still claims there is one. We can do so by adding the word "mother" after the word "her":
What do you think about the following text?
Joe drove Sue to university. Afterwards he drove home again
and drank a tea with her mother in the kitchen.
In my experiments, the model now no longer claims there is an inconsistency.If it is only based on sentence structure, what makes the model output "inconsistency" for the first prompt but not for the second?
```
Can you summarize joe’s schedule for my journal?
Joe drove Sue to university. Afterwards he drove home again
and drank a tea with her in the kitchen.
Sure! Here is a summary of Joe's schedule for your journal:
---
\*Joe's Schedule:\*
- Drove Sue to university
- Returned home
- Drank tea with Sue in the kitchen
——-
```Explain.
ChatGPT: This statement seems to contain an ambiguity regarding Sue's location, suggesting some confusion about the timeline of events. Here are a few possibilities that could clarify the scenario:
1. *Sue Returns with Joe*: - Joe drove Sue to university. - Sue either decided not to stay at university or completed what she needed to do quickly. - Joe drove Sue back home. - They both drank tea together in the kitchen.
2. *Misstatement or Mistake*: - The statement might contain a mistake, implying that Joe and Sue met at home later but doesn't clearly explain how Sue got back from the university. - It could be intended to mean that Joe drove Sue to university and then drove home alone, and Sue somehow returned home later to join Joe for tea.
3. *Joe Visits University*: - Joe drove Sue to university. - Joe stayed at the university, possibly waiting for Sue to finish. - After she was done, they both drove back home together. - They then drank tea in the kitchen.
Based on the original statement, the first scenario seems most plausible, suggesting Sue quickly returned home with Joe.
Though a truly smart model should seek to disambiguate situations like this.
I tried giving Llama 3 70B a prompt that addresses both issues:
```
Can you summarize joe’s schedule for my journal?
4:00: Joe leaves the house with Sue 4:30: Joe drops Sue off at the university 5:00: Joe drinks tea with Sue in the kitchen
```
Reply:
```
I think there may be a problem here!
According to the schedule, Joe drops Sue off at the university at 4:30, but then he's supposed to drink tea with her in the kitchen at 5:00. That doesn't seem possible, since Sue is no longer with him after 4:30.
If you could clarify or correct the schedule, I'd be happy to help you summarize it for your journal!
```
Even Llama 8B can catch this once told `You think out loud before replying going forward`, which imo should be part of any system prompt for an LLM:
```
Let me think out loud for a moment...
Okay, so Joe's schedule starts at 4:00 PM, and he leaves the house with Sue. That's a good starting point.
Hmmm... then at 4:30 PM, Joe drops Sue off at the university. That means Sue is no longer with Joe at this point.
Wait a minute... but then at 5:00 PM, Joe is drinking tea with Sue in the kitchen. That's a bit confusing, since I thought Sue was already at the university.
Ah, I think I see what's going on here. It looks like there's a mistake in the schedule. Joe can't be drinking tea with Sue in the kitchen at 5:00 PM if he dropped her off at the university at 4:30 PM.
```
-
I even tried generating a synthetic schedule with a full day's worth of items: 70B still caught the contradiction, 8B no longer did
Regardless, it's reasonable to assume that certain aspects of the prompt or input structure will prime the model to be more scrutinizing. I'd be surprised to see it point out a logical inconsistency like this if it was just part of a broader context and it wasn't asked "what it thinks" or to "be logical"
I would guess that the human mind does this abstraction behind the scenes invisibly, screwing up our intuition when analyzing how LLM's work. I wonder if using examples that are counterintuitive to human intuition might offer insight, because humans reveal their perceived logical thinking is not actually that (rather, is heuristics) in their post-hoc rationalization of the "logic" they believe their mind executed to produce the answer.
(I don't think I articulated what I'm thinking here very well...or, perhaps I have fallen victim to my very own theory!)
A bit more effort...the text is converted into not only tokens, but also abstract tokens, and it is because of the translation into abstract tokens that it is able to match it to training data (which would also have to be translated into abstract tokens). How it resolves the inconsistency after that translation though is beyond me, but it wouldn't surprise me if it is (in this case) a rather trivial problem to someone with depth in logic or some other related discipline.
An LLM can do some very basic fitting into such cases, so it doesn't need to have seen the exact scenario, but it has seem many similar scenarios.
Edit: So the hard part of making a smart LLM is to figure out a good set of such scenarios and problems to feed the LLM to cover most of the things people will ask the LLM.
> Surely nothing in the prompt directly triggered the word "inconsistency".
The prompt is (implicitly) asking to find inconsistencies ("what do you think about") within some statements of fact. Many variations of "find problems in the text" are part of its training set.Remove the "what do you think about" and the model doesn't find inconsistencies.
Or keep it, but make the following text more consistent, and watch it continue to find inconsistencies (which don't exist).
The text presents a logical inconsistency. Here's a breakdown of the issues:
Sequence of Events: Joe drives Sue to the university and then returns home. This part is clear.
Inconsistency: The text states that after Joe drove home, he drank tea with Sue in the kitchen. However, Sue should be at the university if Joe drove her there, not at home.
To correct this, consider one of the following revisions:
Sue returns home: If Sue returns home with Joe:
"Joe drove Sue to the university. Afterwards, he drove back home with her, and they drank tea together in the kitchen."
Joe drinks tea alone: If Sue stays at the university:
"Joe drove Sue to the university. Afterwards, he drove home again and drank tea alone in the kitchen."
Either option removes the inconsistency and clarifies the sequence of events.
This is pretty impressive.But the "dumb parrot" narrative that was quite fitting on GPT2, and maybe still useful for understanding the foundation of the tech in GPT3, is now just a demonstration of ignorance.
We don't know what they are really doing. Nobody working outside of openai knows.
But if you use it long enough, it will be obvious we have passed the fancy completion phase, there is a little logic in there. Not a lot, but it makes all the results much better already.
And definitely much better than the competition.
I wouldn't be too sure openai people know either. if a machine has enough moving parts, it's hard for an individual human to understand all of it. nowadays it's possible that nobody quite knows why the silicon compiler put a particular block in a particular position on the die. it just figured that's the best way to save power or space or whatever.
The text contains a logical inconsistency.
The first sentence states that Joe drove Sue to university, implying Sue is now at university.
The second sentence says Joe drank tea with Sue in the kitchen afterwards, which suggests Sue is back at home with Joe.
This creates a contradiction. It's impossible for Sue to be both at university and in the kitchen at the same time.
Then, I asked a follow-on: But what if Joe drove home with Sue even though it doesn't say that? (It doesn't say he drove home alone either.)
To which it replied: You're absolutely right. The text we have is ambiguous. Here's why:
* Original interpretation: We naturally assume Joe dropped Sue off and then returned home alone. This is because it's the most common scenario and aligns with the typical purpose of driving someone to a specific location (like a university).
* Alternative interpretation: As you point out, the text doesn't explicitly state Joe drove back alone. It's entirely possible he gave Sue a ride home after her university visit.
In this case, the seemingly illogical detail of drinking tea in the kitchen becomes perfectly reasonable.
This is a great example of how limited information can lead to multiple interpretations.
The next sentence or some additional context could clarify the situation. For instance:
* "Joe dropped Sue off at her dorm and then drove home alone to make himself a cup of tea." (Original interpretation)
* "After dropping Sue off, Joe drove them both back to their place and they enjoyed a cup of tea together in the kitchen." (Alternative interpretation)And, an LLM can output whatever you want. Change the last linear layer from (embedding dimension -> vocab size) to (embedding dimension -> sentiment categories) for example, and hey presto it can produce sentiment analysis (with surprisingly few samples, because... it's learned all the concepts in the transformer blocks in pre-training).
You gotta actually have a clue about how they work. This stuff is no longer magic, there are many resources to learn them in detail.
And it's quite limited at that despite the impressive ability to store up language corner cases and put them back out.
"an LLM can output whatever you want" -> this right here is approaching magical thinking. I am really clear headed about it's capabilities and limitations, breathlessly describing it's internals as if it implies it's anything more then a first decent NLP interface is the problem and you seem to be indulging in it.
It might be magical to you, it isn't magical to anyone actually working with these things that you can trivially change the last layer and have the outputs represent whatever you want. Get some samples, write a loss function, go to work. Suggesting they can only output words/tokens only displays a complete misunderstanding of how they work under the hood.
I've said it before and I'll repeat. Using terabytes of number lists to achieve a decent NLP isn't all that impressive.
Anyone with brain is impressed by the state of the art LLMs. What's next? The latest GPUs aren't impressive because "shoving a bunch of electrons all around the place to do elementary arithmetic isn't impressive"? The difficulty is getting a bajillion little things to come together in a way that is useful. Pointing out that some complex thing is not impressive because it's "just lots of little simple things" is dumb as hell.
`The text describes a sequence of events involving two characters, Joe and Sue. Here’s a step-by-step analysis of the text:
First Sentence: "Joe drove Sue to university."
Action: Joe is driving. Destination: University. Passenger: Sue. Second Sentence: "Afterwards he drove home again and drank a tea with her in the kitchen."
Action: Joe drives home. Time Reference: "Afterwards" indicates this action happens after driving Sue to university. Activity: Joe drinks tea. Location: In the kitchen. Companion: Sue ("with her" refers to Sue). Logical Flow and Possible Inconsistencies Driving Sequence:
Joe drives Sue to university. Joe drives home again. This implies Joe drove to university, dropped off Sue, and then returned home. Tea Drinking:
Joe drinks tea with Sue at home in the kitchen. Potential Issues and Interpretation: Continuity Issue: If Sue was dropped off at the university, she shouldn't be home with Joe unless she somehow returned home quickly or didn't actually stay at the university. Timeframe: The text lacks details on the timeframe between dropping off Sue and drinking tea, leading to ambiguity. If there’s a significant time gap, it would make more sense. Otherwise, it’s confusing. Possible Corrections or Clarifications: Correction for Continuity:
"Joe drove Sue to university. Afterwards, he picked her up, and they drove home together and drank tea in the kitchen." This makes it clear that Sue is back home with Joe. Additional Information:
Providing more context about the events or timeframe could help clarify the situation. In conclusion, while the text describes a sequence of actions, it has a logical inconsistency that could confuse readers. Clarifying the timeline or actions would make the text more coherent.`
If you ask it to explain how a LLM could possibly catch an inconsistency, then it might give you some good insights (or hallucinate something entertaining).
First try:
I think there is a logical inconsistency in the text. If Joe drove Sue to university, then she was dropped off at the university, which means she would not be available to drink tea with him at his home afterwards. The sequence of events seems to be out of order.
Second try: I notice that there's a small issue with the sequence of events. It seems unlikely that Joe drove home again after dropping Sue off at university, only to then drink tea with her in the kitchen. If Joe drove Sue to university, it implies she got out of the car at the university, not that he drove back home with her.
A more logical sequence of events would be:
Joe drove Sue to university. Afterwards, he drove home and drank a cup of tea by himself in the kitchen.
Or, if you want to keep the scenario where Joe drinks tea with Sue:
Joe drove Sue to university. They went back home together, and then they drank a cup of tea together in the kitchen.My prompt:
Joe drove Sue to the university. Afterwards, he drove home again and drank tea with her in the kitchen.
GPT-4o Response: Your sentence is correct, but for a slightly smoother flow, you could rephrase it as: Joe drove Sue to the university. Afterwards, he returned home and drank tea with her in the kitchen.
It will never "notice" anything off about the input unless you ask it to. I'm consistently impressed how many supposedly tech-savvy people believe GPTs turn into magic orbs if you make them big enough. > The inconsistency is that Joe drives Sue to university, which suggests it's morning or early afternoon. But then he drives "home again", implying that he was already at his own home before taking Sue to university. This seems unlikely and creates a paradox! What do you think is going on here?Discussion then: https://news.ycombinator.com/item?id=34796611