ChatGPT Explained: A normie's guide to how it works
jonstokes.com
jonstokes.com
Which incidentally would also imply that a lot of N-grams just don't exist in the training data, causing the model to completely halt when someone says something unexpected.
But that's not the case - instead we convert words into a lot fuzzier float vector space - and then we train the network to predict the next fuzzy float vector. Since the space is so vast, to do this it must learn the ability to generalize, that is, to extrapolate or interpolate predictions even in situations where no examples exists. For this purpose, it has quite a few layers of interconnects with billions of weights where it sums and multiplies numbers from the initial vectors, and during training it tries to tweak those numbers in the general direction of making the error of its last predicted word vector smaller.
And since the N-gram length is so long, the data so large, and the number of internal weights is so big, it has the ability to generalize (extrapolate) very complex things.
So this "probability of next word" thing has some misleading implications WRT what the limits of these models are.
The whole point of mentioning relatedness is to show readers that we’re not dealing with finite n-grams.
The model says, "I don't have a good enough path forwards here, I'll just make one up given the next best thing I have and serve it back"?
Maybe this is why Bing is working differently, they've changed the model or the working to just say, "I don't know" when there isn't enough confidence in what it's generated based off what it finds in it's database?
For example, at a certain size, GPT models start to "learn" how to do basic arithmetic (addition), even for numbers with multiple digits they've never encountered before.
It might look like a small thing, what with computers being quite able to do arithmetic at the base level. But this is a language model, so its a bit different. It learns how to add numbers without "carrying the 1" first, then at a certain larger size also learns to carry the 1, then when even larger it learns to do that across multiple digits... So its not just blindly guessing, its learning the rules of the game (and in some cases some quite complex rules) by building a model of the world made of words.
And the model of digits and addition is just one small bit most likely, as the training space doesn't contain much of that - writing about adding numbers is pretty boring after all. The full model must encode rules and knowledge about a variety of complex things to be able to make reliable predictions. It probably also contains true generalizations that humanity hasn't thought of before, as well as specializations of those generalizations that could be immensely useful.
note: I may have extrapolated a bit more than strictly correct from those two bits, but the accuracy being about 50% indicates carrying is the issue. GPT-3 is at around carrying twice, but hasn't generalized full carrying yet.
It probably also contains true generalizations that humanity hasn't thought of before, as well as specializations of those generalizations that could be immensely useful.
I thought this is quite exciting, I actually think this will become a field of research in itself, and that might spur completely novel new industries and field of technology, research etc.
I'm being optimistic here but I hope we're on the cusp of a greater future, rather than a dystopian one brought on by terminators and AI that replaces all work :)
So it hasn't learned the exact rules for arithmetic, it's learned rules for approximating arithmetic to a decent level of accuracy. Similar to how humans can know the approximate result of an equation before doing the actual math (though GPT is way more precise than humans)
At a certain point, there really is no difference between "being really good at mimicking the sound of Japanese" and actually knowing it, because in order to mimic it to a high level you will have to actually know it. "Mimic the sound of Japanese" is the equivalent task here to "predict the next token in this text".
For now it’s slightly better than fake crowd noise you’re hearing in movies, but still frequently just gibberish.
The Touring test allows unlimited topics, time, etc. There are competitions that use rules heavily favor bots where bots have “won” in the past but they aren’t actually preforming the test.
ChatGPT seems amazing at first, but that’s because its flaws are so novel to them. People just aren’t used to looking for them so they can overlook how quickly it completely forgets about previous parts of a conversation etc.
Passing the Turing test does not imply general intelligence but saying that what it outputs is "just gibberish" is obviously just another hyperbole from you.
I don’t have some arbitrary rules for what passes the Turing test, but it’s about the worst case not the best.
https://en.wikipedia.org/wiki/Computing_Machinery_and_Intell...
> the test as described
I hope you realize that the original Turing test is where you have a man and a woman trying to convince an interrogator that they are of the opposite sex. The test is to replace one with a machine and see if the interrogator would decide the wrong sex as often as when there's an actual human playing.
So if we're talking about the actual test, as described, the most basic bots have passed it a long time ago. If we're talking about the standard interpretation (convince the interrogator that the bot is human) it's a derived version that has no intrinsic rules and was not described by Turing.
I don’t specifically object to changing the judge from interrogation to observation of a conversation. But, it should be clear his version doesn’t have all the loopholes the modern interpretation does.
> wrong sex
You can read the original paper it’s clear in his version the goal for the computer is trying to convince someone communicating with them it’s human even though the form is to convince someone they are male. “The game may perhaps be criticised on the ground that the odds are weighted too heavily against the machine. If the man were to try and pretend to be the machine he would clearly make a very poor showing. He would be given away at once by slowness and inaccuracy in arithmetic.” https://redirect.cs.umbc.edu/courses/471/papers/turing.pdf
It’s also clear he’s referring to the spirit of the game not the specific details: “It might be urged that when playing the "imitation game" the best strategy for the machine may possibly be something other than imitation of the behaviour of a man. This may be, but I think it is unlikely that there is any great effect of this kind. In any case there is no intention to investigate here the theory of the game, and it will be assumed that the best strategy is to try to provide answers that would naturally be given by a man.”
He does give a benchmark of 70% accurate after five minutes of questioning, but that wasn’t success just a benchmark.
You've clearly never used ChatGPT to help you build or troubleshoot anything before.
Try it before you knock it.
That said, it’s got a lot of code to copy from. So I may try again if when I suspect I am reinventing the wheel.
If you told me "We have made a natural language parser/processor that is at human level" I'd think that was a huge step forward for computing. Nobody can argue with that.
And it is quite possible that it's a question without any meaningful answer. We basically have a "gut feeling" that it is a qualitative rather than a quantitative difference, but is that really based on hard data, or it's just more comfortable for us to think that way?
In the meantime, the practical question is - how useful is it? What can it do?
What should I google to understand how a word is encoded as a vector and then vector turned back into word(s)?
The simplest way to get word embeddings (without necessarily building a complex GPT like model) is word2vec: https://towardsdatascience.com/creating-word-embeddings-codi... - the principle is similar but the network is smaller.
People often find it difficult to intuit examples from abstract descriptions. BUT, people are great at intuiting abstractions from concrete examples.
You rarely need to explicitly mention abstractions, in informal talk. People's minds are always abstracting.
> If I’m relating the collections {cat} and {at-cay}, that’s a standard “Pig Latin” transformation I can manage with a simple, handwritten rule set.
Or...
"Translating {cat} to {at-cay}, can be managed with one “Pig Latin” rule:
If input is {cat} then output is {at-cay}."
"Translate" is normie and more context specific to the example than "transformation". "Set" is not "normie" (normies say "collection"), and its superfluous for one rule.Concrete, specific, colloquial, shorter, even when less formally correct, all reduce mental friction.
Yeah that’s exactly the way a nOrMiE would find easy to think about it. Duh.
The author probably should dish out that sentence on his grandparents and see how that would work before putting it on the internet and labeling it as “for normies”.
A generative model is just a computer program. The program is incredibly book smart. No matter how incoherent your ramblings it will find the most likely connection between words, string them into (mostly wrong) candidate sentences with different degrees of certainty, it looks in its database where everything you said is compared to the wrong word combinations assigning scores to each and then it adds up all the points scored and it finds THE most likely correct response: "All gore invented the internet"
There, that is all there is to it.
Same goes for its super brief explanation of a core term like "latent space", and it explain probability distributions by invoking atomic structures rather than a more basic stats example like a bell curve of adult human heights. It's definitely aimed at the sort of "normie" that reads Hacker News rather than actual normal people! I liked it though...
No, “translate” still works here, regardless of Pig Latin not being a language. And I agree with the GP that it’s more intuitively understandable word than “transform.”
translation from one real language to another is so complex that it cannot even be represented in a hash table. Notably, the elements of grammar and syntax are not present in this naive form of translation. It is also lacking the element of bridging understanding, the root of the word translate.
I would even posit that many CS words that are more appropriate here: encoding, compiling, cyphering, mutating, obfuscating,
Thus insisting on a word where critical elements of that word are lacking and despite the existence of more precise language is itself a poor translation.
The fact that it is an invented, derivative, isomorphic language doesn't mean it loses any properties of being a language.
--
Being precise in mathematical or technical terms to a non-technical newbie makes no sense. They need to grasp some basic concepts of what its all about first, in terms that will make immediate sense to them.
Only introduce formalisms when the need for them arises. That way, the reader is not only prepared for them, but can understand the motivation for learning them.
Unless the intro is going to go long and deep, as in the reader is going to become a practitioner, overly precise formalism or language may not add anything at all - because they won't be ready to understand the nuance.
One novel step at a time.
It ends up being easier for the reader, gives them a series of rewarding light-bulb experiences, and requires less explanation in then end. Win-win-win.
Depends on age. I’m 35 and we did electron rings in middle school and probability orbitals in high school.
With no good explanation of course. Just expected to memorize. It didn’t make sense to me until college physics
> If input is {cat} then output is {at-cay}."
Even this can be translated further into “human-speak”:
“Move the first bit of the word to the end, and add ‘ay’. Like, ‘cat’ becomes ‘at-cay’.”
Fixed rule, pattern match rule, collections of rules, ...
Gag.
edit- and then he just casually starts talking about collapsing wave functions and hermeneutics. With no explanations, just assuming the reader knows it. (I had to look up both those concepts) So who exactly is he writing to here?
Post train one LLM to please another LLM, that rates the quality of the first model's responses and calls it on any bullshit! (And vice versa.)
On the other hand, diffusion models seem to have replaced GANs for image synthesis, so perhaps there's something I'm missing, or perhaps there's a way to combine both techniques.
Although an adversarial extra-finger detector seems in order! Extra fingers are not creativity. They are taboo!!
But for text, there is a lot of structure around bullshitting, and fortunately, the internet is full of examples of people calling bullshit. As long as the bullshit adversary has to give a strong critique to back up any bullshit call, it should work.
And bullshit judgments can be de-bullshitted too.
This would be the same as what our social circles, work colleagues, families, etc. do for us, so a circular firing squad of bullshit judges, who also give gold stars for quality responses, seems like a natural solution for models too.
I'm not worried about using approaches like this to converge on truth (model A generates an answer to a question, model B attempts to refute model A). Perhaps we could train model A to generate lies that model B can't detect.
I see this a lot on social media, but it simply isn't true in my experience as someone who uses multiple APIs from OpenAI.
Totally absurd question, totally absurd answer.
"Hello! Can you prove how a martingale converges to the zariski measure on a commutative scheme?"
It replied:
"Hello! The notion of a martingale converging to the Zariski measure on a commutative scheme is a highly advanced and abstract mathematical concept. It is not possible to provide a simple proof for this statement, as it requires a deep understanding of several advanced mathematical concepts such as probability theory, algebraic geometry, and commutative algebra."
It then listed definitions for martingale, Zariski measure, and a commutative scheme, and then concluded with:
"However, to fully understand this concept, you would need to study advanced mathematical topics such as algebraic geometry, commutative algebra, and probability theory."
That seems like a useful enough answer with definitions that effectively converges to "I don't know."
Sorry saw your reply too late.
This could end up being another example of knowledge quality going down as people keep going for the easiest was to get answers, like searching for info on topics on Instagram and TikTok.
My usage has been confined to specific technical inquiries and chatgpts benefit has been colossal in saving me the time of sifting through decade(s) of documentation, forum posts, bug reports, kb articles, etc.
I will be doing some testing of this against the base davinci model over the next few weeks.
(Unfortunately I can't provide you an example. The ChatGPT history is not available right now. Wed Mar 8 02:02:14 PM CET 2023)
I think in the future it'll become less problematic or more transparent about hallucinations
The ‘token window’ section does a fantastic job of answering “but how does it know?”. The ‘lobes of probability’ section does a fantastic of answering “but why does it lie?”.
The ‘whole universes of possible meanings’ bit does an okay job of answering “but how does it understand”, however I think that part could be made more explicit. What made it click for me was https://borretti.me/article/and-yet-it-understands - specifically:
“Every pair of token sequences can, in principle, be stored in a lookup table. You could, in principle, have a lookup table so vast any finite conversation with it would be indistinguishable from talking to a human, … But it wouldn’t fit in the entire universe. And there is no compression scheme … that would make it fit. But GPT-3 masses next to nothing at 800GiB.
“How is it so small, and yet capable of so much? Because it is forgetting irrelevant details. There is another term for this: abstraction. It is forming concepts. There comes a point in the performance to model size curve where the simpler hypothesis has to be that the model really does understand what it is saying, and we have clearly passed it.”
If I was trying to explain that to normies, I would try to hijack the popular “autocomplete on steroids” refrain. Currently it seems like normies know “autocomplete for words”, and think when you put it on steroids you get “autocomplete for paragraphs”. Explain to them that actually, what you get is “autocomplete for meanings”.
(Please feel free to use these ideas in your post if you like them, don’t even think about crediting me, I just want to see the water level rise!)
https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
As far as I understand, in this article he argues that in the, say, ChatGPT output, compression happens.
But does it really?
To make a similarly low resolution metaphor, a “bayesian kaleidoscope” of a language model doesn’t necessarily mean it blurs the “word pixels” it is moving around. Because moving them around, rearranging them is what it essentially does, even if in opaque ways; but not degrading them, not changing letters in words or deliberately algorithmically messing up the word order in a sentence.
To make sense of the “image” an LLM produces is left up to us, and therefore, it is also up to us to decide whether any compression of anything has happened. And then, how do you measure it?
If you cut a painting into pieces, then glue them back together at random, thus making a new painting, would that constitute a “compression” or just a new painting, which could be worse or better than the original?
I quite like Chiang’s writing, but not this time. If anything, his take on this undermines what he previously wrote a little bit, painting him more of an LLM that he probably would like to admit :)
That isn't a perfect metaphor, but it explains very well how it can do most of the things it can do. The lossy compression means that it can work with large prompts and just capture their essence instead of trying to look them up literally, and the lossy decompression lets it vary its output and the text will move in slightly different directions instead of just repeating text it has seen. The magical bit is that this compression and decompression is much smarter than before, it parses text to a format much closer to its meaning than before, and that lets us do the above much more intelligently.
Edit: Thinking a bit, maybe you could make these model way cheaper to run if we would make them work as a compression to meaning rather than the huge models they are now? They do have internal understanding/meaning of the tokens it gets, so it should be possible to create a compression/decompression function based on these models that transforms text into its world model state, and then once we start working with world model states things should be super cheap relative to what we have now.
Also maybe it doesn't have lossy decompression and get words with similar meaning, but that is another way I see the models could be smaller and cheaper while keeping their essence. The Markov chain step could be all it uses currently. But it definitely creates that space and Markov chain, because it parses the previous thousand or so tokens and uses those to guess the next token, that is a Markov chain. It just has a very sophisticated way of parsing those thousand tokens into a logical format.
I dislike that interpretation. It suggests it builds a very basic statistical model, but a very basic statistical model simply wouldn't be able to do what these models can do.
Or alternatively, if you want to consider the model as a markov chain mapping the probability from the previous four thousand tokens to the next token then the space is astronomically large. Beyond astronomically and even economically large, there are ~50,000^4096 possible input states.
Why do you think that? Why do you think a basic statistical continuation of the logic of a text wouldn't do what the current model does? There are trillions of conversations out there it can rely on to continue the text, people playing theatre, people roleplaying, tutorials, people playing opposite games, people brainstorming etc. Create a parser that can parse those down to logic, then make a markov chain based on that, and I have no problem seeing the current ChatGPT skills manifesting from that.
> Or alternatively, if you want to consider the model as a markov chain mapping the probability from the previous four thousand tokens to the next token then the space is astronomically large. Beyond astronomically and even economically large, there are ~50,000^4096 possible input states.
Yes, that is the novel thing, it compresses the states down to something manageable without losing the essence of the text, and then builds a model there of likely next token.
> We all learn in basic chemistry that orbitals are just regions of space where an electron is likely to be at any moment
You may be surprised.
Latent space is then introduced in a side note:
> “latent space is the multidimensional space of all likely word sequences the model might output”
So the simple overview is that "Oh hey, it's just like electron orbitals - you know except in a multidimensional space of word sequences"?
The end part is probably the most useful, describing how these things work in a bit more practical sense. Overall this feels like it introduces the fact the model is static and it has a token window in a very complicated way.
I understand the process abstractly, but I am unable to grok the details about how it's able to take my vague interpretation of what I want and then write code and actually give me exactly what I wanted.
Trial and error in terms of tweaking the system prompt is surprisingly instructive.
But with ChatGPT, it just seems like magic.
Then once you have such a parser you can now make a logical Markov chain, it predicts the continuation of your text based on continuations of things that looks similar in the logic space rather than the text space, and that is what you get back from the model.
So ChatGPT does what we could have easily done if we had if natural language was more logical. Then we could make a simple model based on all the logic encoded in all writing of humanity that just looks up human done logic similar to what you asked for and then return an average of those logics. Now since you want human language output it now has to translate that back to human text from logic space, and it could be done in English, Spanish, it could speak like a pirate etc.
As a heuristic, I see descriptions falling into simple buckets:
- stories that talk about tokens
- stories that don’t talk about tokens
Anything discussing technical details such as tokens never seems to really get around to the emergent properties that are the crux of ChatGPTs importance of society. It’s like talking about humanity by describing the function of cells. Accurate, but incomplete.
On the other hand, higher-level takes happily discuss the potential implications of the emergent behaviours but err on the side of attributing magic to the process.
I haven’t read much, to be fair, but I don’t see anyone tying those topics together very well.
I think that is due to LLMs being somewhat magical. I think that the wolfram article "What Is ChatGPT Doing … and Why Does It Work?" captures this beautifully.
Attention is a recurrence relationship that gets gradually pruned.
What does the author think knowing actually is if not a convergence of probability distributions?
I'd like for all these "let me explain to you how things work" type write ups to first define what their understanding of the key preexisting concepts are (like what it is to know something). So when they try to explain the new concepts I don't waste my time learning concepts formed on a flawed understanding of their foundational counterparts.
1. Hello, what can you tell me about the Legendre polynomials, angular momentum, and the three-dimensional visualization of atomic orbitals in the context of quantum mechanics?
2. So, can the three-dimensional representations of atomic orbitals seen in chemistry textbooks be considered to be solutions of the Schrodinger differential equations in the context of Legendre polynomials?
3. What's the nature - mathematically speaking - of the radial and angular combination of the Legendre polynomials and the hydrogenic radial functions that give rise to the textbook probablility density visualizations of atomic orbitals?
4. Tell me about Hartree-Fock representations of the molecular orbitals of the helium atom?
5. Compare and contrast the Hartree-Fock method and the more recently developed Density Functional Theory approach, please, in the context of simple molecular orbital theory.
That's pretty impressive, probably a better and more concise explanation than a lot of professionals in the field could give off the top of their heads. However, it's also interesting to see where ChatGPT is remarkable untrustworthy - Linux terminal command line options, for example, it often gets those wrong (disastrously so in some cases, happily I double-checked first). It also has problems with modular arithmetic... This one always delivers, it seems, and has been doing so for a month at least:
Q: So is seven congruent to four, in the context of modulus three?
> "No, seven is not congruent to four modulo three. To determine whether two integers are congruent modulo a given modulus, we need to compute the difference between them and check if it is divisible by the modulus.
In this case, we have:
7 - 4 = 3
Since 3 is not divisible by 3, we conclude that 7 is not congruent to 4 modulo 3.
We can also verify this by checking the remainders when 7 and 4 are divided by 3:
7 ÷ 3 = 2, remainder 1 4 ÷ 3 = 1, remainder 1
Since the remainders are not the same, we can again conclude that 7 is not congruent to 4 modulo 3."
Hmmm.... this seems wrong...
At the point it’s generating the next word, it knows what its preceding words were. With a conceptual representation of various rhyme schemes, subsequent words will (probably) fit that form.
In my experience, when it can't come up with a good rhyme, it tends to fib it. But also keep in mind that, as it operates on token representation of words internally, its notion of what rhymes and what does not is far from perfect. This is more noticeable in languages other than English.
Super flawed explanation but for non-compsci friends it has helped them understand the mechanism a BIT better.
I tried to offer that it is pretty cool, but it's just software that basically presents search engine results in a different manner along with a few other tricks, but it's not "HAL".
I live in a very red and rural area so that probably has something to do with it. They love to have new things to complain about that have no effect on any of us at all.
That seems an ungenerous interpretation. I don't doubt that their understanding is full of science-fiction inspired fear, but the implications and dangers of this tech is a hotly debated topic among informed experts. So, even if their specific fears are ungrounded, their fear may not be (ie, effects on the economy, culture, education, especially as the tech advances, which it will).
over the past three years the entire world was impacted by a dire health crisis where misinformation played a large role in distorting public perception. this has direct impacts on public health (people not wearing masks, refusing vaccines) and has spillover effect into other parts of people's lives (political polarization around said issues).
what you see as a pretty cool toy could also easily be abused as a giant round the clock fake news generator. it doesn't matter if the text is true or even makes any sense... an alarming amount of people will take anything they read as fact without investigating the sources. this can be done as is with chatgpt right now. it has obscenity filters sure, but fake news is trying to pass as legitimate reporting, so it will probably be framed in a tone that escapes the obvious filters they have.
then consider the implications for robotexting, phishing, automated bots that pretend to be you to customer service chats, social media bots, messaging app scammers. all of these things are currently problems that can cause harm both personal and societal... and chatgpt will make it easier and cheaper to scale them up to new levels.
I increasingly hear things that suggest that those who wore masks and got vaccinated were the ones who were actually misinformed. Of course, you won't hear any of that on CNN
The only folks who were misinformed are those who never took the time to learn about masks and vaccinations. But to be fair that is a huge number of folks here in the U.S.
I've yet hear anyone talk about someone they infected who was killed by it though. With over 1.1 Million dead from it here in the U.S. that's astonishing. So is the fact that it is still killing around 2000+ a week here.
This seems like an unnecessary invocation of negative stereotypes. What makes you think people outside outside that demographic don't have similar thoughts? Anecdata and all that.
The fact that people think they are rationally debating issues is the problem.
Same could be said about Covid, vaccines, and masking during a viral pandemic.
My guess is it's an addiction to the outrage FOX News sells (and others like it).
What bugs me most about ChatGPT is that it doesn't attribute where it got the results it offers. For example, I asked it: "how to save a document with PouchDB" and it showed me the code, and while I didn't compare it, it looks like a copy and paste from PouchDB.com, and that should be referenced with a link in the response if that's the case, but it was not.