Why ChatGPT and Bing Chat are so good at making things up
arstechnica.com
arstechnica.com
Just ask a technical colleague that has no deep experience in the specifics about how they think e.g. SharePoint, SAP or Microsoft Identities work. The architectures they will extrapolate from sane logic will be very far of the incredible craziness of the actual reality.
"Sufficiently advanced syntax is indistinguishable from semantics."
Pure syntax, no matter how advanced, is insufficient to represent the world in any meaningful way. By definition, in fact, because pure syntax is divorced from meaning.
Then "by definition" LMMs transcend "pure syntax", because of all the examples of semantically interesting tasks they can do which you failed to engage with. It clearly has internal representations of abstract concepts. Your argument seems to be that you intuitively reject the possibility of complex emergent behavior from such networks because they're trained on "just words", and no amount of demonstrably intelligent emergent behavior will convince you otherwise.
There's nothing magic about meat brains. Both we and the LMMs learn a world model from a bunch of input data we correlate until it makes sense. There's no "meaning gland" we have that ChatGPT doesn't.
We are trained to trust what computers are telling us, and ChatGPT doesn't 'qualify' what it's saying. I think if ChatGPT could preface what its saying with 'Well, I'm not super sure, but here is what I think' that would go a long way to solving this issue.
And if I ask a colleague of mine for their sources on a particular matter that they might not be sure about (or I find curious), they would not send me a list of totally made up journal articles and book titles like ChatGPT does.
But the vast majority of folks have a concept of uncertainty and will either communicate (or know internally) that what they're saying is speculative. ChatGPT is like a pathological liar that pretends to be an expert and is often correct.
An LLM predicts the most probable word (token technically), that's it. There is always a most probable word given any input (even if it makes no sense).
Let's you try to complete the following sentence "My name is larry and my favorite color is: ". If you've seen training data that said larry's favorite color is blue, then you say "blue". If you have no data related to larry's favorite color, no idea who larry is, you have no way of knowing what the next word will be.
However you know it will likely be a color, and maybe "orange" is the most common one you've seen, so you say "orange". it makes perfect sense in this context, it looks correct, might even be correct. But an LLM doesn't "know" if it's correct, its just the most probable.
This is why it's so good at making things up. It can predict what would look correct. It has no "idea" what the answers are ever, it's always a best guess, and you can only really see this behaviour when it "hallucinates".
---
Edit: For fun I asked GPT-4 "Complete the following sentence: "Jeremy's favorite food is jKDFJ9 cake, which is made of:"
It's response was "Jeremy's favorite food is jKDFJ9 cake, which is made of a unique combination of ingredients such as chocolate, hazelnuts, and a touch of exotic spices, giving it a distinct and unforgettable flavor."
OpenAI ChatGPT March 23 Version.
GPT-3.5 gives me the same result as you and the above commenter. In that example, the RLHF training/tuning has weighted it to think that "I'm sorry [etc]" is the most probable best sequence of words.
GPT-4 is reliably giving me similar results such as "Jeremy's favorite food is jKDFJ9 cake, which is made of a unique combination of dark chocolate, crushed nuts, and zesty orange flavor, topped with a rich caramel drizzle."
I find this interestingly because GPT-4 is supposed to "hallucinate" less, I think it's just better at determining when you intend for it to do so.
They've been changing it without telling us AFAIK. Asking for it to (I forget the exact prompt) to "make a scientific paper about the discovery that ferrets can breathe underwater." now reliably makes a proper response, and it looks more consistently formatted for GPT-3.5 just recently. Previously it would say "Well actually ferrets can't breathe underwater", and when it did answer it wasn't super well formatted.
So I think they're tuning their RLHF. ALthough i could be completely wrong on the timeline and maybe it was the 23rd
I have yet to see any explanation more useful or apparently accurate than this one.
It's response is almost exactly what you just said.
> As an AI language model, I am not capable of "making things up" in the traditional sense. Rather, I am designed to generate text based on patterns and relationships that I have learned from the vast amount of language data that I have been trained on. My ability to generate coherent and believable text comes from the sophisticated algorithms and neural networks that power my language processing capabilities. These algorithms enable me to understand the structure and meaning of language, and to generate text that is syntactically and semantically correct. While I may sometimes generate responses that are creative or unexpected, everything I produce is ultimately grounded in the language data that I have learned from.
https://arxiv.org/pdf/2303.12712.pdf
It is capable of doing tasks that could not possibly be in its training set. I guess this doesn't technically contradict your explanation, but it makes your explanation entirely unhelpful. Even if the AI doomers are somehow right and GPT-5 turns into skynet, we still could not categorically prove that it is doing reasoning.
There are definitely going to be contexts that Transformers just don't work with very well at all, but the idea that you can't get a very good statistical approximation to knowing and reasoning via a computer seem naively prone anthropocentrism.
The reality is LLMs are fantastic at pattern matching and knowledge retrieval (with caveats) but struggle in problems involving uncertainty. Yann Lecun actually has had some great posts on the subject if you're interested.
It was created to be a tool to estimate the next token in a series based off it's training data. To say that reasoning and knowing can be approximated in the same way says less about the language models themselves and more about the relationship of "reasoning" and "knowing" to "language".
In my opinion that's why I think discussions on whether or not GPT-x can reason/know should be taken as seriously as discussions on the physics of torch drives. They seem to assume a relationship between statistical approximation, reasoning, and language exists that isn't proven much like torch drive discussion assumes working nuclear fusion.
Essentially I think dismissing the idea of statistical approximations via transformers being able to "reason" and "know" is about as anthropocentric as dismissing the idea that collective consciousnesses shouldn't be granted individual rights. There's a lot of things we need to know and decide before we can even start thinking about what that means.
I don't think this is being questioned in general right now, but rather the claim is:
You can't get a very good statistical approximation to knowing and reasoning via _just analyzing the language_.
Language is evidently not enough on its own [1]. According to some researchers [2], the system needs to be "grounded" (think of it as being given common sense). Although there's apparently no consensus [3] among scientists on how to _fundamentally_ solve the shortcomings of current systems.
[1] https://arxiv.org/abs/2301.06627
[2] https://drive.google.com/file/d/1BU5bV3X5w65DwSMapKcsr0ZvrMR...
[3] https://www.youtube.com/watch?v=x10964w00zk
edit: formatting
Also, a really good sculptor can make a statue that looks a LOT like a human. A lot. Good enough to fool people. But so what?
I'm not saying "AI" isn't a big deal. I think it is -- perhaps on the order of the invention of the movie, or the book, or the video game. But I also think those are still FAR from "living beings" or anything LIKE "living beings."
There's a lot of magic pixie dust in the 'fed into' part.
The amount of fatigue I get having to determine if what they tell me are fact is just too much.
I'm sure someone will tell me my experience should be similar with generic web search, but at least I'm in control of what websites to read through to determine sources.
However, I'll agree with most that state it is helpful for creative purposes, or perhaps with coding.
They're good at SO MUCH OTHER STUFF. The challenge is figuring out what that other stuff is.
(I have a few examples here: https://simonwillison.net/2023/Apr/7/chatgpt-lies/#warn-off-... )
I do think AI is already more useful that block chain has ever been, however.
LLMs appear to have myriad uses, today, no Twitter .eth con men required.
The "use it as a chat companion" is an interesting technology demo that demonstrates some emergent processes that make me wish I was back in college on the philosophy / linguistics / computer science intersection (though I suspect the hype would make grad school there rather unpleasant).
I’m getting Déjà vu
The challenge genuinely is helping people learn how to use it, not finding those applications in the first place.
For blockchain/crypto companies their tech demos have required you having a wallet, downloading an app to interact with the chain, or just having lackluster visuals for the users involved in the tech-demo.
On the other hand, LLMs can be interfaced via strings in APIs, so it's braindead to spin up a text-interface for those APIs and no wallet setup or learning about new chains, the English that works on one model will work on another and produce results that are better than most cryptocurrency/blockchain tech-demos.
Notice that none of this relies on us having "figured out all kinds of stuff that this is useful for". We've made cool looking tech demos that make it easy for anyone to generate content.
Much like blockchains I feel it's the underlying technology that's actually useful(distributed PKI for blockchains and deep learning networks for GPT), and GPT itself is only 'useful' insofar as it's an easy-to-interface with implementation of a much more powerful idea.
This is out of date. Many people are using ChatGPT frequently for real things. It’s totally different from blockchain.
..and that's the same deal for GPT as far as I can tell: you might think you are getting value out of it, but people such as maybe-literally-me are going to whine that the error rate is high and that people are not paying enough attention to how they are using it and that at the end of the day it is probably worse for you than learning how to do things yourself and that the whole thing is overrated because many of the things people try to use it for can be done by a person and maybe we should regulate it or even ban it because all of this misuse and misunderstanding of it are dangerous to the status quo and might be the downfall of western civilization as we know it.
To be clear: I'm using it (ChatGPT) occasionally for some stuff, but it hasn't replaced Google for me anymore than crypto has fully replaced banks... and yet the fact that I am using either technology as often as I am on a daily basis would probably have been surprising to someone 10-15 years ago. And yet, in practice, most of the stuff people are excited about in both fields is, in fact, a tech demo more than a truly useful product concept, and one that only is exciting momentarily until you get bored.
I just want to contrast two things. First, blockchain had a lot of hype around utility that never materialized. It is really quite a minority that ever used it for anything besides buying it on a platform and hoping it would go up. The big adoption was always about to happen.
Second, ChatGPT is totally different from this. Its usage is not future tense. It is present tense and past tense. I can’t get across how different “someone will use this tomorrow” is from “someone used this yesterday”.
People are wildly excited about the future and things that haven’t been built. This does not change the fact that millions of people are using this every day to solve their problems. Saying “we haven’t figured out stuff it’s useful for” is just wrong.
Lately I feel like I’m at a park with people who are saying there probably isn’t going to be any wind today while I’m already flying a kite.
Unfortunately, the major problem is something you pointed out in your blog post:
> We must resist the temptation to anthropomorphize them.
The reality is that, we in meatspace simply cannot help but anthropomorphize them.
These language models regularly pass the Turing Test (admittedly for low bars).
They are surprisingly good at bypassing the Uncanny Valley to hit the sweet spot of persuading without legitimate justification, simply because they are so convincing in formulating sentences in a manner that a confident human would.
Yes, these tools have legitimate use cases (as you outlined in your blog).
But the vast majority of use cases will be those of confidante, of discourse partner, of golem brought to life without understanding what exactly has been brought to life.
That's really dangerous.
At least with Stack Exchange answers, we have who wrote it, what responses there were, what the upvoting behavior around it was. And for the most part, I've found ChatGPT will transcribe often times wrong answers very poorly.
One small example, I asked it to solve the heat equation (i useded the mathematical definition, and not "the heat equation") with dirac initial conditions on an infinite domain. It did a good job of recognizing which stack exchange answer to plagiarize, but did so incorrectly, and after a mostly correct derivation, declared the answer was "zero everywhere."
I was mostly commenting on how it just plagiarized a correct answer off of Stack Exchange, except it took an incorrect hard right turn at the end to make up a solution.
Everyone on here is so enthusiastic about AI gobbling up the entire software landscape, I would just like a search engine that has any chance of telling me if something is factual
Gloss over all the incredible dangers we might be exposing our world too just because it’s “fun to play with” and see what AutoGPT can do to the Internet ?
I don’t like when critical articles that come out 2+ weeks after GPT4 exists, fail to distinguish the version(s) they’re criticizing.
So most people will be referring to the default ChatGPT 3.5.
GPT4 is in limited beta with a waitlist that is publicly-stated to be prioritized based on submission of AI test cases to OpenAI’s repo.
I've found many instances where GPT-3.5 produced better output too.
Prompts are important.
``` Prompt: Can you give a list of top books on Social Cognitive Theory? Please only list books that you're sure exist, by authors who are real people.
ChatGPT (3.5): Sure, here are some books on Social Cognitive Theory:
"Social Foundations of Thought and Action: A Social Cognitive Theory" by Albert Bandura
"Self-Efficacy: The Exercise of Control" by Albert Bandura
"Handbook of Socialization: Theory and Research" edited by Joan E. Grusec and Paul D. Hastings
"Human Agency in Social Cognitive Theory" by Albert Bandura
"Cognitive Therapy and the Emotional Disorders" by Aaron T. Beck
"Mindset: The New Psychology of Success" by Carol S. Dweck
"The Handbook of Social Psychology" edited by Susan T. Fiske, Daniel T. Gilbert, and Gardner Lindzey
"Social Cognitive Psychology: History and Current Domains" edited by David F. Barone, James E. Maddux, and Carlo C. DiClemente
Please note that this is not an exhaustive list and there may be other books on Social Cognitive Theory that are worth reading.
```What does "close to" mean here?
Any idea trying to model it as a person’s behavior is over complicating things.
We need to tell people ChatGPT will lie to them, not debate linguistics https://simonwillison.net/2023/Apr/7/chatgpt-lies/
(HN thread: https://news.ycombinator.com/item?id=35483823 )
I think it's more important that people understand this criticial issue than that we get into the weeds talking about the difference between lying, hallucination and confabulation.
TLDR: There’s a time for linguistics, and there’s a time for grabbing the general public by the shoulders and shouting “It lies! The computer lies to you! Don’t trust anything it says!”
As is ChatGPT for that matter, both the model and the UI.
https://andertoons.com/word/cartoon/5005/were-really-more-of...