We Found an Neuron in GPT-2
clementneo.com
clementneo.com
It's crazy that large language models work so well just by being trained as a next-word-prediction model over a large amount of text data. We know how image models learn extract the features of an image through convolution[1], but how and what LLMs learn exactly remain a black box. When we dig deeper into the mechanisms that drive LLMs, we might get closer to understanding why they work so well in some senses, and why they could be catastrophic in other cases (see: the past month of search-based developments).
I find trying to understand and reverse-engineer LLMs to be a personally exciting endeavour. As LLMs get better in the near future, I sure hope our understanding of them can keep up as well!
We know the architecture of LLMs because we created it, but we don't yet have the same level of understanding about them, or the same quality of analytical tools for reasoning about them.
The point of using a CNN instead of a FCN is that you force it to learn in a certain way that prevents overfitting. But given a sufficient dataset, and proper data augmentation you would expect a FCN to be able to identify objects regardless of translation. It's just that a CNN would train easier and better, with a smaller network (a FCN doing convolutions would be very wasteful).
That's why traditionally you would pick your architecture to help it learn in a certain way (images=cnn, text=rnn/lstm/gru). But the nice thing about transformers is that they are more general.
In the case of CNN the reason it works is that an image of an object X is still an image of object X if the X is shifted left or right. The property is translationally invariant. CNN are basically the simplest way to encode translational invariance.
That's the geometric deep learning theory, isn't it? Do you know if there's a list somewhere of exactly what invariance has which ways to simulate it? Like an overview?
So you can’t extract the weightings for ‘an’ discretely because those weightings encode its connection with all the other words and combinations and sequences or clusters of words it might ever be used with, including their weightings with other preceding words, and their relationships, etc, etc.
Come to think of it, when someone teaches me a new concept, the principle of mass conservation, for instance, in some sense they are transferring their embedding into my brain, further on I will relate to mass conservation through what that person taught me. The transfer is a very lossy process, sure, but a transfer with reintegration nonetheless. Perhaps "mortal computation" [4] is a requirement.
[1] https://en.wikipedia.org/wiki/Grandmother_cell
[2] https://www.youtube.com/playlist?list=PL8FnQMH2k7jzPrxqdYufo...
[3] https://www.youtube.com/watch?v=kTcRRaXV-fg
[4] Geoffrey Hinton, The Forward-Forward Algorithm: Some Preliminary Investigations, chapter 8, https://www.cs.toronto.edu/~hinton/FFA13.pdf
Firstly even if there is such a cell that only fires for one face, or perhaps also the person’s name, it doesn’t mean there aren’t other cells that fire for that person, or for people in general including that person. Without those as well, that neurons responses might not mean anything to the rest if the brain. It’s a thought experiment but never really demonstrated.
Also even if this is true in the very strongest sense. Say there is one neuron that uniquely and discretely fires in response to thinking about that one person. What defines a neuron isn’t just its internal behaviour. It’s also the pattern of inputs that influence it, and the pattern of outputs it sends out. It’s the connections and dependencies on the weightings and signals and responses from all the cells it’s connected to. Including the specific unique ways all those neurons are connected, or not connected to all the other cells in the brain. It’s al, the specifics of that connectedness that are what makes the behaviour of that neuron meaningful.
If you took that neuron and implanted it into another brain, you’d need to hook it up to the neurons in that brain such that it gets exactly the same stimuli, in the same order, with the same strength, every time it needs to fire. The same applies to its output, all the neurons it’s connected to would have to interpret its firing behaviour in the exact same way the other neurons in the original brain did. But there’s no guarantee any of those connected mechanisms work or are physically connected in the same way, or even a vaguely similar or compatible way in the new brain.
This is Wharton professor Ethan Mollick playing with the new Bing chat, which seems considerably more advanced than ChatGPT (based on GPT-4 perhaps?).
Here he asks it to write something using Kurt Vonnegut's rules of writing.
https://twitter.com/emollick/status/1626084142239649792
It seems hard to explain how Bing/GPT could have generated the Vonnegut-inspired cake story, having ingested the rules, without planning the whole thing before generating the first word.
It seems there's an awful lot more going on internally in these models than a mere word by word autoregressive generation. It seems the prompt (in this case including Vonnegut's rules) is ingested and creates a complex internal state that is then responsible for the coherency and content of the output. The fact that it necessarily has to generate the output one word at a time seems to be a bit misleading in terms of understanding when the actual "output prediction" takes place.
Note too that despite the output being sampled from a distribution based on a "randomness" temperature, there are many case where what it is trying to say so much constrains the output that certain words/synonyms/concepts are all but forced.
1. we don't know what they(coding layer between bing and GPT) look up and store as a prompt aka working memory.
2. it can do the equivalent of receiving it's own prompt silently.
I seen with code it outputs the step for the code then writes the code.
so there's some kind of plan and execute going on. maybe it can do that in model some how
The simple answer is that the internal state that picks the next token is stable over iterations so that the model can follow a consistent plan over multiple token outputs. Then as the plan "unfolds" in the output tokens, these tokens help stabilize the plan further, thus creating consistency over long generations.
https://arxiv.org/abs/2202.05262
Locating and Editing Factual Associations in GPT
See also new interesting developments breaking the connection between "Locating" and "Editing":
https://arxiv.org/abs/2301.04213
Does Localization Inform Editing? Surprising Differences in Causality-Based Localization vs. Knowledge Editing in Language Models
Sometimes if someone is working though a complex thought and they're not really sure where they're going, they'll pause while thinking of the word they want to use, and might sound like
"the discussion is an... an... epistemological one"
Obviously they may have been conscious that the next word was going to start with "epi..." and they are just trying to remember the word, but I think sometimes they really don't consciously know what they're going to say, only unconsciously.
It reminds me of a recent New Yorker article about how people think, where the author realized they often have no idea what they're about to say before they open their mouths. [1]
1. https://www.newyorker.com/magazine/2023/01/16/how-should-we-...
Disclaimer: Uneducated opinion on my behalf, I'm a hobbyist only.
I hear that in the voice of Agent Smith, from The Matrix.
I certainly don't know what words a sentence is going to end with when I'm thinking or saying the first words in the sentence. I just think or say the sentence from start to finish, never knowing what the next word is going to be as I'm thinking the current one, and by the end of it I've thought or said a full sentence that makes sense.
It makes me feel like I’m not very good at conversation though, especially smalltalk.
I think there's something like a stage design in the brain:
- symbolic or model deliberation
- verbal expression
- vocalization
And each of those stages can be consciously introspected on, but people will naturally develop more or less ability to introspect on it. I think when people say they "don't mentally verbalize", what is actually happening is they just haven't happened to develop conscious introspection of the verbal expression stage. But I'd expect that this can be trained.
(Conversely, sometimes people introspect so much on verbal expression that it becomes an inherent part of the way they think. Brains are weird and wonderful!)
Part of this auto-pilot is learning to recognize if you have the answer or not. I will not launch into a sentence without a strong feeling I understand the topic and know what to say. I just don't need to prepare the exact words in order to do it. I still consciously "check" that I know, but that just requires a boolean answer instead of a word by word crafted sentence.
Like
[Conceptual stage] --??-- consciousness
V
[Verbalization stage] --??-- consciousness
V
[Vocalization]
So if you don't have conscious access to verbalization, you only realize how a thought "will sound" after you say it. Conversely, if you don't have conscious access to conceptualization, you end up thinking that "thinking" always involves "thinking out loud", because "thinking out loud" (verbalizing) is the only way you have to query your conceptual layer. You literally only become aware of your own thinking after the thought is already pretty far along. At the extreme, you can have conscious access to neither stage and require vocalization to reflect on your own thinking.
> Conversely, if you don't have conscious access to conceptualization, you end up thinking that "thinking" always involves "thinking out loud", because "thinking out loud" (verbalizing) is the only way you have to query your conceptual layer. You literally only become aware of your own thinking after the thought is already pretty far along
What I'm saying is that even if you do have conscious access to verbalisation stage, the more you force yourself to let it be handled subconsciously, the better at communicating quickly and effectively. This is what I meant by autopilot.
Spending active conversation time on conscious verbalization seems to me as inefficient as verbalising words and "speaking them out" as you read a book, usually known as subvocalization.
I tend to think of consciousness as the "debug mode" of the brain.
I’m thinking training on movies/tv has the physical element that pure text is missing.
I think of it a bit like walking. You can think about it, focus on it, control it as you please, but most of the time you just do it without thinking.
Consciousness is the caller. You/It can consciously manually call `moveLeftFootUp()`, then `moveRightFootDown()`. Or maybe you were calling `walk(speed="normal")` and stepped in and started debugging that function's code at that level, step by step. Also, these functions sometimes raise exceptions, which are either handled by the function's caller automatically or bringing it to the caller's attention (i.e. the `conciousness()` main loop).
Learning to walk involves first manually calling `moveLeftFootUp()` and `moveRighFootDown()` order (once you have drafted those functions) in different order to get right how that should be done, then prototyping some `walk()` function code. The initial version of the `walk()` function at the begining isn't very robust and doesn't handle a lot of edge cases, thus raising exceptions all the time and requiring a lot of conscious effort. Of course, you are also adjusting `moveRighFootDown()` and `moveLeftFootUp()` at the same time or maybe creating `moveFoot(feet,direction)` function, etc.
But in the end, after all the fine adjustments of the code, you basically get the code for `walk()` right, it stops raising exception's to the main loop and doesn't require too much effort. You can just call the `walk()` function and it just works automatically (unless you step in with the debugger) - or you can continue creating new functions that call `walk()` inside those, confidently.
"Every time I fire a linguist, the performance of the speech recognizer goes up". - Frederick Jelinek
"It would be interesting to know how different a model would be if it operated on eg dependency trees instead of the linear list of tokens." - indeed, this is obviously interesting, so people have tried that a lot for many models, but IMHO it's probably now almost decade since the consensus is that in general end-to-end training (once we became able to do it) work better than adding explicit stages in between, e.g. for any random task I would expect that doing text->syntax tree->outcome is going to get worse results than text->outcome, because even if the task really needs syntax, the syntax representation that a stack of large transformer layers learns implicitly tends to be a better representation of the natural language than any human linguist devised formal grammar, which inevitably has to mangle the actual languge to fit into neat human-analyzable 'boxes'/classes/types in which it doesn't really fit and all the fuzzy edge cases stick out. Once you remove the constraint that the grammar must be simplified enough for a human to be able to rationally analyze and understand, and (perhaps even more importantly?) abandoning the need to prematurely disambiguate the utterance to a single syntax tree instead of accepting that parts of it are ambiguous (but not equally likely), processing works better.
It's just one more reminder of the bitter lesson (http://incompleteideas.net/IncIdeas/BitterLesson.html) which we don't want to accept.
So there seems to be such emergent mechanisms in the model that have arisen because of the end-to-end training, which we don't exactly understand yet.
First off, we know that overall the concept of language is humans is an emergent phenomenon. It developed from natural selection from simple components so there's validity that the same thing can occur in an LLM where some overarching emergent structure develops from simple primitives.
At the same time we do know that a sort of universal grammar exists among humans. Our language capacity is biased in a certain way and that it is unlikely for it to learn languages of a very extreme and divergent grammar from the universal one discovered by Noam chomsky. That means our brain is unlikely to be as universally simple as an LLM.
I think the key here is that the human mind has explicit linguistic tools but the these tools are still emergent in nature.
Has it changed since then? Have they found a significant number of human civilizations that use divergent grammars?
Harris’s operator grammar was based on set theory however Chomsky was enamored with formal logic and ran in that direction. Also, Chomsky became famous while Harris didn’t.
Operator grammar is self-discoverable and he published a full description of English grammar in the 1980’s using this theory and an extension in the 1990’s which generalized to other languages.
It is not fully deterministic (final word selection and ordering is probabilistic), but it is much more convincing to me and in line with how we understand brains to work.
Over the last 20 years Chomsky has begun to fall out of favor because it’s just so complicated and requires external structures, etc.
Any theory that explains all the facts is valid, there can be more than one competing theory. Fame of the author absolutely matters and so does elegance of the theory.
Chomsky has held on so long because it did pretty well and he became very famous, but over time it has needed larger and larger patches to cover its flaws, so other explanations are gaining traction once again.
No the elegance of the theory does not matter. The truth of the theory does. I'm saying does the universal grammar you describe, is it actually universal? Amongst all the languages in the world, do they all share this other grammar you describe? If not then the theory is incorrect. It doesn't matter how elegant the theory is.
I brought this up specifically because your reasoning wasn't "scientific" it was more along the lines of logical elegance.
>Chomsky has held on so long because it did pretty well and he became very famous, but over time it has needed larger and larger patches to cover its flaws, so other explanations are gaining traction once again.
See this makes sense and is basically the answer to the question I am asking. So you're basically saying that there were flaws? As in he had to make up new rules constantly because the science was contradicting his theory. Am I correct in this characterization?
I’m not an expert, but if I recall I think the big flaw was insistence on determinism.
And to my knowledge operator grammar is universal to the degree there isn’t a counter example in the dozen of languages he explored and hundreds he had others help him with. But to my knowledge the only language he fully mapped was English.
Edit: take a look at the article on Wikipedia.
https://arxiv.org/abs/1905.05950
I think it's a mistake to discount the psychology of learning. It's like saying that "calories-in, calories-out" is all there is to weight-loss. Strictly true, but not helpful for 90% of people.
Probability and observation are all that is required to understand a language.
I don’t think this is entirely fair. Generative grammars have been produced for a huge variety of non-European languages, even non-Indo-European languages, and can account for tremendous diversity in linguistic rules. Even languages without fixed word orders or highly synthetic languages can be represented.
Linguistics isn’t focused on the problem of outputting reasonable-sounding text responses. Instead, it seeks to transparently explain how language works and is structured, something that GPT does not do.
>grammar as we know it was devised for the Latin language and linguists spend most of the time attempting to fit other languages into neat boxes that the Latin grammar wasn't designed for....Chomsky tried to . All solve this problem with Universal Grammar.
Well first of all, every language has some form of grammar. If there were no rules that made meaning depend on word order, there would be nothing different between what the cat ate and what ate the cat. To claim Grammar is some Latin/Western system being unduly applied is absurd.
You are correct UG was invented as a to explain something about English. Howevert was not an attempt to solve "this problem ("this problem" being conforming to expected systems of Grmmar.) The problem he was interested in solving was about the ability to learn language rapidly despite not having many negative examples of how not to talk. Pay attention here researchers, the lesson maybe extends to you too soon.
UG also isn't a specific enough thing you can check against in a directory. Its a theory, namely a theory that if you study language features you'll eventually discver some things are invaiant. Using it directly and immediately based on nothing but Chomsky's conjecture...all I can focus on passing out at screen goodnightn
I'm not sure I understand why this is an open question. While I get that GPT-2 is predicting only one word at a time, it doesn't seem that surprising that there might be cases where there is a dominant bigram (ie "an apple" in the case of their example prompt) that would trigger an "an" prediction, without actually predicting the following word first.
Am I missing something?
(Admittedly, all this ML/AI stuff is still beyond my current level of understanding, so I'm sure my thinking here is off.)
You could probably test this by seeing if prompts containing a lot of nouns that start with a vowel sound results in output that contains a higher proportion of otherwise unrelated first-vowel nouns. (ie your prompt includes lots of apples, apricots, avocados, asparagus, aubergines, elderberries, eggplants, endives, oranges, olives, okras, onions and you count the proportion of non-food nouns in the result that start with a vowel and non-vowel sound).
To rephrase that for this case: what is the specific mechanism in GPT-2 that (1) makes it realise that the word 'apple' is significant in this prompt, and (2) use that knowledge to push the model to predict 'an'? Finding this neuron would only answer the some portion of (2).
(And to rephrase this for the general case, which gives us the initial question: How does GPT-2 know when, given a suitable context, to predict 'an' over 'a'?)
But GPT can’t think ahead what token it will add after the one it is on. Or can it? It could “predict” internally the word apple for the next “meaningful” word and output ‘an’ because of this.
Exactly what you mean by bigram.
But yeah, it should be easier when you're just dealing with text I guess
une hache - ma hache
une horloge - mon horlogeObviously any romance language includes gender, so you need to use the correct gendered article before the noun.
Japanese has different counting words depending on what you're counting.
I don't know if Mandarin has anything similar.
But then: “ Testing the neuron on a larger dataset”
If I follow correctly, they test a bunch of different completions that contain “an”. So they are not just detecting the bigram “an apple”, but the common activation among a bunch of “an X” activations where X is completely different.
Of course it can only answer the next word, because there is only room in its outputs for the next word. But it has to compute much more. It a huge hidden internal state where it has to first encode what the given sentence is about, then predict some general concept in which the continuation goes, decide the locally correct syntactical structure and only from this you can predict the next word.
1. It's kinda interesting because this is a clear case where the model must be thinking beyond the next token, whereas in most contexts it's hard to say whether the model thinks ahead at all (although I would guess that it does most of the time).
2. More importantly, the key question here is how it works. We're not surprised that it has this behavior, but we want to understand which exact weights and biases in the network are responsible.
Note also that this is just the introductory sentence and the rest of the article would read exactly the same without it.
> it doesn't seem that surprising that there might be cases where there is a dominant bigram [...] that would trigger an "an" prediction, without actually predicting the following word first
btw I don't really understand what you mean by this. Bigrams can explain the second prediction in a two word pair but not the first.
Its answer?
"Red apple".
...well played.
Rephrase "The model is good at picking the correct article for the word it wants to output next" to "After having picked a specific article, the model is good at picking a follow-up noun that matches the chosen article". Nothing about the second statement seems like an unlikely feat for a model that only predicts one word at a time without any thinking ahead about specific words.
>I climbed up the pear tree and picked a pear. I climbed up the apple tree and picked
The argument made in the article (IMO an extremely convincing one) is that it wouldn't be able to predict the word 'an' except by observing that the word afterwards must be apple. Otherwise why not pick 'a'?
I don't see how you get to this conclusion. From all the training data it has seen, "an" is the most probable next word after "I climbed up the tree and picked up". The network does not need to know anything about the apple at this point. Then, the next word is "apple" (with an even higher probability I guess).
So (again, I am very much not an expert so please correct me if I'm wrong) I guess my analogy would be if instead of predicting one word at a time, it predicted one letter at a time. At some point, would only be one word that could fit. So if prompted with "SENTE", and it returned "N", that doesn't mean that it's thinking ahead to the "CE" / knows that it is spelling "SENTENCE" already.
Is that a correct way to think of it?
1: https://www.cell.com/cancer-cell/pdf/S1535-6108(02)00133-2.p...
input text -> input tokens -> input embeddings -> model -> output embeddings -> output tokens -> output text
Tokens aren't necessarily words: they can be fragments of words and you can check out this behavior here: https://platform.openai.com/tokenizerFor instance, "an eagle" is tokenized to [an][ eagle], but "anoxic" is tokenized to [an][oxic], so just looking for the [an] token is not sufficient. Therefore, you would need to map the output text all the way back into the model to figure out what neuron(s) in the model would generate "an" over "a". Since the bulk of GPT is all unsupervised learning, any connections it makes in its neural network is all emergent.
Yeah except it doesn't necessarily take the strongest one, you can set a 'temperature' parameter which determines how often you should pick the second choice, etc.
>How is it "clocked" to get successive tokens?
You add each new token to the input prompt to generate the next one.
OMFG so it really is a super advanced auto-complete!!
Yeah, this is not AGI or anywhere close. That explains how it picks up context, and also how it can lose the plot after a bit due to limited input size.
Ie this seems like such a cool area to be in but the data volumes required are huge, complex, etc. Code is simple, cheap, lean, etc by comparison.
Do we have any insight on how this area of research could be usable with less hardware and data? Is there a visible future where a guy and a laptop can make a big program? (without depending on tech getting small/cheap in 50 years or w/e)
The pile-10k dataset they used for analysis is 33MB, and GPT2 runs ok on a CPU. For the full 10K analysis it's probably quicker to get a GPU though.
This is correct. We did almost all this work on a Macbook Pro. Although for the pile-10k dataset analysis we used an A100 GPU because it would take many hours to run the whole thing through GPT-2 on a laptop.
we find that there is no distinction between individual high level units and random linear combinations of high level units, according to various methods of unit analysis. It suggests that it is the space, rather than the individual units, that contains the semantic information in the high layers of neural networks.
[1] Intriguing properties of neural networks
It could have been hallucinating, of course. But it does seem like it occasionally alters already generated words as it goes, at least I think I've seen it do that.
The slow progression of ChatGPT output is just a property of the output layer. The language engine doesn't work slowly like that, and once a token has been generated it can't backtrack.
In other words, it outputs one token at a time.
My take on why they have build the output layer like it is, is that next to feeling more human, it also forces you to be a bit more thoughtfull with your requests, and thus spam the system less. In the end it is still really expensive to run these models..
No, it's literally a for loop in python that runs the whole thing from scratch[1] after appending each new token. No artificial slowdowns, what you're seeing is literally what it's spitting out in real time
Here's an example
https://github.com/karpathy/minGPT/blob/master/mingpt/model....
[1] Some things can be cached such as kv, but still
Or an aitch / h, but only sometimes.
That's the opposite of how it works (as you demonstrated in your own comment, "not a rule...")
The comment you responded to was correct, or at least very reasonable, and I'm really not sure which part you disagree with.
There's a few words in English that begin with the letter h but not with the consonant sound /h/ (depending on your accent!), mostly because they're French in origin: historic is one, and there's also hour, heir, honour. "It is an honour to be here"
It's silly to act like there's a right or wrong. Dialects and accents are a special thing. Just be consistent.
Solution get rid of article. Someone want to know if thing is definitive thing, can figure it out from context!
(also, I'm in the opposite superstate to the one you are in -- both "a SQL parser" and "an SQL parser" look fine to me, although I internally pronounce the first as "a seekwul parser" and the second as "an ess cue ell parser")
a sequel parser is also fine *
* my old Unix colleagues will remember that Sequel is a brand of database that nobody else remembers anymore so tended to use the initialism.
We found a Neuron in a Neural Network
Nothing new to see here. They pinpoint the nodes where the training bumped up the numbers for one token while not firing other tokens.Yes, I’m being a bit reductionist, but GPT -> transformer architecture -> neural network. We just have more detailed techniques, a lot more data, storage, and processing power now. But the basics of how a NN works hasn’t changed.
Not one neuron which was more relevant than all the others put together.
It kinda confirms that it's all just turtles, all the way down.