Language models can explain neurons in language models
openai.com
openai.com
"... our technique works poorly for larger models, possibly because later layers are harder to explain."
And even for GPT-2, which is what they used for the paper:
"... the vast majority of our explanations score poorly ..."
Which is to say, we still have no clue as to what's going on inside GPT-4 or even GPT-3, which I think is the question many want an answer to. This may be the first step towards that, but as they also note, the technique is already very computationally intensive, and the focus on individual neurons as a function of input means that they can't "reverse engineer" larger structures composed of multiple neurons nor a neuron that has multiple roles; I would expect the former in particular to be much more common in larger models, which is perhaps why they're harder to analyze in this manner.
I wonder how often this happens in the universe...
Think of business / artistic / cultural leaders nurturing protégés despite not totally understanding why they’re successful.
Of course those protégés have agency and drive, so maybe not a perfect analogy. But I’m going to stand by the point intuitively even if a better example escapes me.
Is that evident already or are we fitting the definition of intelligence without being aware?
In human brains, language is only a way to communicate thoughts in concept form, though we also seem to use language to communicate abstract thoughts to ourselves to break them apart/down in a way (imo).
I'd love to see someone train a model on the level of GPT4 to generate abstract thoughts/ideas based on input/context and then pair this model with GPT4 co-operatively and continue to train, such that the flow of abstract ideas is parsed by GPT. But like...how do you even train a model that operates on abstract ideas, there doesn't seem to be any way to do this.
An example: we are not very good at creating flight, the one birds do and humans always regarded as flight, and yet we fly across half the globe in one day.
Going up three meters and landing on a branch is a different matter.
https://pub.towardsai.net/ais-mind-reading-revolution-how-gp...
so why not have them decode sequential dense vectors of their own activations?
As for the majority scoring poorly, they suggest that most neurons won't have clear activation semantics so that is intrinsic to the task and you'd have to move to "decoding the semantics of neurons that fire as a group"
The more interesting question is why are intelligence/beauty/consciousness emergent properties that exist in our minds.
Perhaps intelligence is like a black box input to our bodies (call it the "soul", even though this isn't testable and therefore not a hypothesis). The mind therefore wouldn't play any more of a role in intelligence than the eye. And I'm not sure people would say the eye is necessary for understanding intelligence.
Now, I'm not really in a position to argue for such a thing, even if I believe it, but I'm curious what argument you might have against it.
1. Neurons connect all our senses and all our muscles.
2. Neurons are the definitive difference between the brain and the rest of the body. There is “other stuff” in the brain, but it’s not so different from the “other stuff” that’s in your rear end.
Don’t underestimate what a neuron can do. A single artificial neuron can fit a logistic regression model. A quarter of a million is on then scale of some our our largest AI, and biological neurons are far more connected than ANN. An ant quite likely has a more powerful brain than GPT-4.
Our digestive systems appear to be important to our behaviour, though. Some recent work in mice showed that if colonised with bacteria from faeces of humans with autism, the mice would begin to show autistic behaviours.
So, not sure your argument here is especially strong.
Ants have developed architecture, with plumbing, ventilation, nurseries for rearing the young, and paved thoroughfares. Ants practice agriculture, including animal husbandry. Ants have social stratification that differs from but is comparable to that of human cultures, with division of labor into worker, soldier, and other specialties that do not have a clear human analogy.
Ants enslave other ants. Ants interactively teach other ants, something few other animals do, among them humans. Ants have built "supercolonies" dwarfing any human city, stretching over 5,000 km in one place. And ants too have a complex culture of sorts, including rich languages based on pheromones.
Despite the radically different nature of our two civilizations, it is undeniable from an objective standpoint that this level of society has been achieved by ants.
[0]: https://www.reddit.com/r/unpopularopinion/comments/t2h1vs/an...
Humans so far have done a great job at destroying nature faster than any other kind could.
And GPT4 was created for profit.
The paper is somewhat new so I haven't done a proper review to know if it's solid work yet, but this may offer some context for some of the comments in this thread.
Natural language understanding comprises a wide range of diverse tasks such as textual entailment, question answering, semantic similarity assessment, and document classification. Although large unlabeled text corpora are abundant, labeled data for learning these specific tasks is scarce, making it challenging for discriminatively trained models to perform adequately. We demonstrate that large gains on these tasks can be realized by generative pre-training of a language model on a diverse corpus of unlabeled text, followed by discriminative fine-tuning on each specific task. In contrast to previous approaches, we make use of task-aware input transformations during fine-tuning to achieve effective transfer while requiring minimal changes to the model architecture. We demonstrate the effectiveness of our approach on a wide range of benchmarks for natural language understanding. Our general task-agnostic model outperforms discriminatively trained models that use architectures specifically crafted for each task, significantly improving upon the state of the art in 9 out of the 12 tasks studied. For instance, we achieve absolute improvements of 8.9% on commonsense reasoning (Stories Cloze Test), 5.7% on question answering (RACE), and 1.5% on textual entailment (MultiNLI).
Should that be measured in number of nuclear power plants needed to run the computation? Or like, fractions of a small star’s output?
Exactly. Especially:
> ...the technique is already very computationally intensive, and the focus on individual neurons as a function of input means that they can't "reverse engineer" larger structures composed of multiple neurons nor a neuron that has multiple roles;
This paper just brings us no closer to explainability in black box neural networks and is just another excuse piece by OpenAI to try to please the explainability situation that has been missing for decades in neural networks.
It is also the reason why they cannot be trusted in the most serious of applications which such decision making requires lots of transparency rather than a model regurgitating nonsense confidently.
Like say, in court to detect if someone is lying? Or at an airport to detect drugs?
But, yes, those are also good examples of what we shouldn't be doing, but are going to do anyway.
Evolution is still doing it's thing.
This seems like a novel approach to try to tackle the scale of the problem. Just because the earliest results aren’t great doesn’t mean it’s not a fruitful path to travel.
Is this true? I thought explainability for things like DNNs for vision made pretty good progress in the last decade.
Doesn't this criticism also apply to people to some extent? We don't know what the purpose of individual brain neurons is.
The machine model is not only a black box, but one incapable of understanding anything about its input, "thought process", or output. It will blindly spit out a response based on its training data and weights, without knowing the difference whether it true or false, meaningful or complete gibberish.
In some cases, you can clearly see neurons that specialize to different areas of the function being modeled, like this one: https://i.ameo.link/b0p.png
This OpenAI research seems to be feeding lots of varied input text into the models they're examining and keeping track of the activations of different neurons along the way. Another method I remember seeing used in the past involves using an optimizer to generate inputs that maximally activate particular neurons in vision models[2].
I'm sure that's much more difficult or even impossible for transformers which operate on sequences of tokens/embeddings rather than single static input vectors, but maybe there's a way to generate input embeddings and then use some method to convert them back into tokens.
[2] https://www.tensorflow.org/tutorials/generative/deepdream
I'd be curious to see Softmax Linear Units [1] integrated into the possible activation functions since they seem to improve interpretability.
PS: I share your curiosity with respect to things like deep dream. My brief summary of this paper is that you can use GPT4 to summarize what's similar about a set of highlighted words in context which is clever but doesn't fundamentally inform much that we didn't already know about how these models work. I wonder if there's some diffusion based approach that could be used to diffuse from noise in the residual stream towards a maximized activation at a particular point.
On first look this is genius but it seems pretty tautological in a way. How do we know if the explainer is good?... Kinda leads to thinking about who watches the watchers...
The paper explains this in detail, but here is a summary: an explanation is good if you can recover actual neuron behavior from the explanation. They ask GPT-4 to guess neuron activation given an explanation and an input (the paper includes the full prompt used). And then they calculate correlation of actual neuron activation and simulated neuron activation.
They discuss two issues with this methodology. First, explanations are ultimately for humans, so using GPT-4 to simulate humans, while necessary in practice, may cause divergence. They guard against this by asking humans whether they agree with the explanation, and showing that humans agree more with an explanation that scores high in correlation.
Second, correlation is an imperfect measure of how faithfully neuron behavior is reproduced. To guard against this, they run the neural network with activation of the neuron replaced with simulated activation, and show that the neural network output is closer (measured in Jensen-Shannon divergence) if correlation is higher.
To be clear, this is only neuron activation strength for text inputs. We aren't doing any mechanistic modeling of whether our explanation of what the neuron does predicts any role the neuron might play within the internals of the network, despite most neurons likely having a role that can only be succinctly summarized in relation to the rest of the network.
It seems very easy to end up with explanations that correlate well with a neuron, but do not actually meaningfully explain what the neuron is doing.
The reliability question is of course the main issue. If you don't know how the system works, you can't assign a trust value to anything it comes up with, even if it seems like what it comes up with makes sense.
It seems NN output could be trusted in scenarios where a test exists. For example: "ChatGPT design a house using [APP] and make sure the compiled plans comply with structural/electrical/design/etc codes for area [X]".
But how is any information that isn't testable trusted? I'm open to the idea ChatGPT is as credible as experts in the dismal sciences given that information cannot be proven or falsified and legitimacy is assigned by stringing together words that "makes sense".
I understand that around the 1980s-ish, the dream was that people could express knowledge in something like Prolog, including the test-case, which can then be deterministically evaluated. This does really work, but surprisingly many things cannot be represented in terms of “facts” which really limits its applicability.
I didn’t opt for Prolog electives in school (I did Haskell instead) so I honestly don’t know why so many “things” are unrepresentable as “facts”.
"Answer this question in the form of a testable prolog program"
The bigger value here in the near-term is _explicability_ rather than alignment per-se. Potentially having good explicability might provide insights into the design and architecture of LLMs in general, and that in-turn may enable better design of alignment-schemes.
I really like their approach and I think it’s valuable. And in this particular case, they do have a way to score the explainer model. And I think it could be very valuable for various AI Safety issues.
However, I don’t yet see how it can help with the potentially biggest danger where a super intelligent AGI is created that is not aligned with humans. The newly created AGI might be 10x more intelligent than the explainer model. To such an extent that the explainer model is not capable of understanding any tactics deployed by the super intelligent AGI. The same way ants are most probably not capable of explaining the tactics delloyed by humans, even if we gave them a 100 years to figure it out.
https://openaipublic.blob.core.windows.net/neuron-explainer/...
"Suddenly, DM-sliding seems positively whimsical"
https://openaipublic.blob.core.windows.net/neuron-explainer/...
https://www.thecut.com/2016/01/19th-century-men-were-awful-a...
Aww, that's so nice of them to let the community do the work they can use for free. I might even forget that most of OpenAI is closed source.
"In one well-known experiment, a split-brain patient’s left hemisphere was shown a picture of a chicken claw and his right hemisphere was shown a picture of a snow scene. The patient was asked to point to a card that was associated with the picture he just saw. With his left hand (controlled by his right hemisphere) he selected a shovel, which matched the snow scene. With his right hand (controlled by his left hemisphere) he selected a chicken, which matched the chicken claw. Next, the experimenter asked the patient why he selected each item. One would expect the speaking left hemisphere to explain why it chose the chicken but not why it chose the shovel, since the left hemisphere did not have access to information about the snow scene. Instead, the patient’s speaking left hemisphere replied, “Oh, that’s simple. The chicken claw goes with the chicken and you need a shovel to clean out the chicken shed”" [1]. Also [2] has an interesting hypothesis on split-brains: not two agents, but two streams of perception.
[1] 2014, "Divergent hemispheric reasoning strategies: reducing uncertainty versus resolving inconsistency", https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522
[2] 2017, "The Split-Brain phenomenon revisited: A single conscious agent with split perception", https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf
Even if you accept classic theory (e.g. hemispheric localization and the homunculus) which most experts don't all this suggests is that the brain tries to make sense of the information it has and in sparse environments it fills in.
How does this make our behavior "mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning" as most humans don't have a severed corpus callosum.
The discussion starts with:
"In a healthy human brain, these divergent hemispheric tendencies complement each other and create a balanced and flexible reasoning system. Working in unison, the left and right hemispheres can create inferences that have explanatory power and both internal and external consistency."
But the bottom line is that introspection is not necessarily reliable.
The existence of cognitive dissonance suggested in your citation is in no way analogous to "our own explanations about ourselves and our behaviour are mostly lies, fabrications, hallucinations, faulty re-memorization, post hoc reasoning" and in fact supports the opposite.
https://www.health.harvard.edu/blog/right-brainleft-brain-ri... :
> But, the evidence discounting the left/right brain concept is accumulating. According to a 2013 study from the University of Utah, brain scans demonstrate that activity is similar on both sides of the brain regardless of one's personality.
> They looked at the brain scans of more than 1,000 young people between the ages of 7 and 29 and divided different areas of the brain into 7,000 regions to determine whether one side of the brain was more active or connected than the other side. No evidence of "sidedness" was found. The authors concluded that the notion of some people being more left-brained or right-brained is more a figure of speech than an anatomically accurate description.
Here's wikipedia on the topic: "Lateralization of brain function" https://en.wikipedia.org/wiki/Lateralization_of_brain_functi...
Furthermore, "Neuropsychoanalysis" https://en.wikipedia.org/wiki/Neuropsychoanalysis
Neuropsychology: https://en.wikipedia.org/wiki/Neuropsychology
Personality psychology > ~Biophysiological: https://en.wikipedia.org/wiki/Personality_psychology
MBTI > Criticism: https://en.wikipedia.org/wiki/Myers%E2%80%93Briggs_Type_Indi...
Connectome: https://en.wikipedia.org/wiki/Connectome
The post you are replying to is talking about the small subset of individuals who have had their corpus callosum surgically severed, which makes it much more difficult for the brain to send messages between hemispheres. These patients exhibit “split brain” behavior that is well studied by experiments and can shed light into human consciousness and rationality.
I think the correct statement is "so far the answer is we don't know"
After breaking my arm, split in two, pinching the nerve and making me unable to move it for about a year, I still feel as if the arm is "someone else's", as if I am moving an object in VR, not something which is "me" or "mine".
Reading/listening to someone like Robert Sapolsky [1] makes me laugh I could have ever hallucinated about such a muddy, not even wrong concept as "free will".
Furthermore, between the brain and, say, the liver there is only a difference of speed/data integrity inasmuch as one cares to look for information processing as basal cognition: neurons firing in the brain, voltage-gated ion channels and gap junctions controlling bioelectrical gradients in the liver, and almost everywhere in the body. Why does only the brain has a "feels like" sensation? The liver may have one as well, but the brain being an autarchic dictator perhaps suppresses the feeling of the liver, it certainly abstracts away the thousands of highly specialized decisions the liver takes each second solving adequately the complex problem space of blood processing. Perhaps Thomas Nagel shouldn't have asked "What Is It Like to Be a Bat?" [2] but what is it like to be a liver.
[1] "Robert Sapolsky: Justice and morality in the absence of free will", https://www.youtube.com/watch?v=nhvAAvwS-UA
[2] https://en.wikipedia.org/wiki/What_Is_It_Like_to_Be_a_Bat%3F
And that’s just the polar opposite of having a meaningful will at all. It is good that you are pretty much deterministic. You shouldn’t be deciding meaningful things randomly. If you made 20 copies of yourself and asked them to support or oppose some essential and important political question (about human rights, or war, or what-have-you) they should all come down on the same side. What kind of a Will would that be that chose randomly?
* in a safe setting with support, of course.
I am morbidly curious how people are going to creatively explain away the more challenging insights AI gives us in to what consciousness is.
https://www.mpg.de/research/unconscious-decisions-in-the-bra...
The idea that there should not be any neural activity before a conscious decision is straight-up dualism---the intangible soul makes a decision and neural activity follows it to carry out the decision.
An alternative way of understanding that result is that the neural activity that precedes the "conscious decision" is the brain's mechanism of coming up with that decision. The "conscious mind" is the result of neural activity, right?
if it turns out that true, it’s truly amazing how well we convince ourselves that we’re in control.
but if our brain controls our actions and not our consciousness, then what is the purpose of consciousness?
There is no "their" and there is no "thought process" . There is something that produces text that appears to humans like there is something like thought going on (cf the Eliza Effect), but we must be wary of this anthropomorphising language.
There is no self reflection, but if you ask an LLM program how "it" knows something it will produce some text.
To be clear, you're saying that we should just dismiss out-of-hand any possibility that an LM AI might actually be able to explain its reasoning step-by-step?
I find it kind of charming actually how so many humans are just so darn sure that they have their own special kind of cognition that could never be replicated. Not even with 175,000,000,000 calculations for every word generated.
All this talk of AGI and sentience and so on is premature and totally unfounded . It's pure sci fi, for now at least.
Above you said about AI LMs:
> There is no "their" and there is no "thought process"
So, unless you're claiming that humans lack a thought process as well, then you're arguing that humans are special.
> All this talk of AGI and sentience and so on is premature and totally unfounded
I don't see any mention of AGI or sentience in this thread?
Also, I don't think anyone could read this transcript with GPT-4 and still claim that it's incapable of a significant degree of self-reflection and metacognition:
And when Hinton says at MIT, "I find it very hard to believe that they don't have semantics when they consult problems like you know how I paint the rooms how I get all the rooms in my house to be painted white in two years time," I believe he's commenting on the ability of LLM's to think on some level.
1. Show GPT-4 a GPT-produced text with the activation level of a specific neuron at the time it was producing that part of the text highlighted. They then ask GPT-4 for an explanation of what the neuron is doing.
Text: "...mathematics is done _properly_, it...if it's done _right_. (Take ..."
GPT produces "words and phrases related to performing actions correctly or properly".
2. Based on the explanation, get GPT to guess how strong the neuron activates on a new text.
"Assuming that the neuron activates on words and phrases related to performing actions correctly or properly. GPT-4 guesses how strongly the neuron responds at each token: '...Boot. When done _correctly_, "Secure...'"
3. Compare those predictions to the actual activations of the neuron on the text to generate a score.
So there is no introspection going on.
They say, "We applied our method to all MLP neurons in GPT-2 XL [out of 1.5B?]. We found over 1,000 neurons with explanations that scored at least 0.8, meaning that according to GPT-4 they account for most of the neuron's top-activating behavior." But they also mention, "However, we found that both GPT-4-based and human contractor explanations still score poorly in absolute terms. When looking at neurons, we also found the typical neuron appeared quite polysemantic."
I think LLMs are "Semantic Clouds of Words" + grammar and syntax generator. Someone could just discard the grammar and syntax generator, just use the semantic cloud and create the grammar and syntax by himself.
For example, in writing a legal document, a slightly educated person on the subject, could just use the relevant words put into an empty paper, fill in the blanks of syntax and grammar, alongside with the human reasoning which is far superior than any machine reasoning, till today at least.
The process of editing the GPT* generated documents to fix reasoning is not a negligible task anyway. Sam Altman mentioned that: "the machine has some kind of reasoning", not a human reasoning ability by any means.
My point is, that LLMs are two programs fused into one, "word clouds" and "syntax and grammar", sprinkled with some kind of poor reasoning. Their word clouding ability, is so unbelievable stronger than any human it fills me with awe every time i use it. Everything else is, just whatever!
Looking at it this way, I honestly wouldn't be surprised if that's exactly how "System 1" (to borrow a term from Kahneman) in our brains works.
What I'm saying is:
> In my opinion, in case there is a way to extract "Semantic Clouds of Words", i.e given a particular topic, navigate semantic clouds word by word, find some close neighbours of that word, jump to a neighbour of that word and so on, then LLMs might not seem that big of a deal.
It may be much more of a deal than we'd naively think - it seems to me that a lot of what we'd consider "thinking" and "reasoning" can be effectively implemented as proximity search in a high-dimensional enough vector space. In that case, such extracted "Semantic Cloud of Words" may turn out to represent the very structure of reasoning as humans do it - structure implicitly encoded in all the text that was used as training data for the LLMs.
Yes, exactly that. That's what GPT4 is doing, over billions of parameters, and many layers stacked on top of one another.
Let me give you one more tangible example. Suppose Stable Diffusion had two steps of generating images with humans in it. One step, is taking as input an SVG file, with some simple lines which describe the human anatomy, with body position, joints, dots as eyes etc. Something very simple xkcd style. From then on, it generates the full human which corresponds to exactly the input SVG.
Instead of SD being a single model, it could be multimodal, and it should work a lot better in that respect. Every image generator suffers from that problem, human anatomy is very difficult to get right.[1] The same way GPT4 could function as well. Being multimodal instead of a single model, with the two steps discreet from one another.
So, in some use cases, we could generate some semantic clouds, and generate syntax and grammar as a second step. And if we don't care that much about perfect syntax and grammar, we feed it to GPT2, which is much cheaper to run, and much faster. When i used the paid service of GPT3, back in 2020, the Ada model, was the worst one, but it was the cheapest and fastest. And it was fast. I mean instantaneous.
>the very structure of reasoning as humans do it
I don't agree that the machine reasons even close to a human as of today. It will get better of course over time. However in some not so frequent cases, it comes close. Some times, it seems like it, but only superficially i would argue. Upon closer inspection the machine spits out non sense.
[1] Human anatomy, is very difficult to get right, like an artist. Many/all of the artists, point out the fact, that A.I. art doesn't have soul in the pictures. I share the same sentiment.
What is known is that these internal thoughts get erased each time a new token is generated. That is, it's starting from scratch from the contents of the text each time it generates a word. But you could postulate that similar prompt text leads to similar "thoughts" and/or navigation of the concept web, and therefore the thoughts are continuous in a sense.
But todays networks lacks the recursion(feedback where the output can go directly to the input) that is needed for the type of internalized thoughts that humans have. I guess this is one thing you are pointing at by mentioning the continuousnes of the internals of LLMs.
When someone states definitively what LLMs can or cannot do, that is when you know to immediately disregard them as the waffling of uninformed laymen lacking the necessary knowledge foundations (cognitive/neuroscience/philosophy) to even appreciate the uncertainty and finer points under discussion (all the open questions regarding human cognition etc).
They don't know what they don't know and make unfounded assertions as result.
Many would do to refrain from speaking so surely about matters they know nothing about, but that is the internet for you.
What if you ask it to synthesize multiple internal streams of thought, for an ensemble of interior monologues, then have all those argue with each other using logic and then present a high level answer from that panoply of answers?
To me it makes no sense to say that a LLM could explain its own reasoning if it does no (logical) reasoning at all. It might be able to explain how the neural network calculates its results. But there are no logical reasoning steps in there that could be explained, are there?
IANAE but although an LLM meets the definition of a Markov Chain as I understand it (current state in, probabilities of next states out), the big black box that spits out the probabilities could be doing anything.
Is it fundamentally impossible for reasoning to be an emergent property of an LLM, in a similar way to a brain? They can certainly do a good impression of logical reasoning- better than some humans in some cases?
Just because an LLM can be described as a Markov Chain doesn’t mean it _uses_ Markov Chains? An LLM is very different to the normal examples of Markov Chains I’m familiar with.
Or am I missing something?
In any case, coemu is an interesting related idea to constrain AIs to thinking in ways we can understand better:
https://futureoflife.org/podcast/connor-leahy-on-agi-and-cog...
https://www.alignmentforum.org/posts/ngEvKav9w57XrGQnb/cogni...
There's no formal axiom system being dealt with here, afaict?
Do you just generally mean "there may be some kind of self-reference, which may lead to some kind of liar-paradox-related issues"?
Well, what exactly would we be showing that these models can’t do? Quines exist, so there’s no general principle preventing reflection in general. We can certainly write poems (etc.) which describe their own composition. A computer can store specifications (and circuit diagrams, chip designs, etc.) for all its parts, and interactively describe how they all work.
If we are just saying “ML models can’t solve the halting problem”, then ok, duh. If we want to say “they don’t prove their own consistency” then also duh, they aren’t formal systems in a sense where “are they consistent (as a formal system)?” even makes sense as a question.
I don’t see a reason why either Gödel or Turing’s results would be any obstacle for some mechanism modeling/describing how it works. They do pose limits on how well they can describe “what they will do” in a sense of like, “what will it ‘eventually’ do, on any arbitrary topic”. But as for something describing how it itself works, there appears to be no issue.
If the task to give it was something like “is there any input which you could be given which would result in an output such that P(input,output)” for arbitrary P, then yeah I would expect such diagonalization problems to pop-up.
But a system having a kind of introspection about how it works, rather than answering arbitrary questions about its final outputs (such as, program output, or whether a statement has a proof), seems totally fine.
Side note: One funny thing: (aiui) it is theoretically possible for a oracle that can have random behavior, to act (in a certain sense) as a halting-oracle for Turing machines with access to the same oracle.
That’s not to say that we can irl construct such a thing, as we can’t even make a halting oracle for normal Turing machines. But, if you add in some random behavior for the oracles, you can kinda evade the problems that come from the diagonalization.
Some training forms include entailment : “if A then B”. I hope this is first order logic which does have an axiom system :)
Hofstadter talks about something similar in his books.
Of course, if you were trying to use GPT-4 to explain GPT-4 then I think the Gödel incompleteness theorem would be more relevant, and even then I'm not so sure.
In a real sense, all of the future discoveries of mathematics already exist in the "training set" of our present understanding, we just haven't thought it all the way through yet. If we discover something new, can we say that the concept didn't exist, or that it "couldn't be inferred" from previous work?
I think the same would apply to LLMs and their understanding of the way we encode information using language. Given their radically different approach to understanding the same medium, they are well poised to both confirm many things we understand intuitively as well as expose the shortcomings of our human-centric model of understanding.
EDIT: I see below you gave some examples, like invention of language before it existed, and new theorems in math that presumably would be of interest to mathematicians. Those ones are fair enough in my opinion. The AI isn't quite good enough for those ones I think, but I also think newer versions trained with only more CPU/GPU and more parameters and more data could be 'AI scientists' that will make these kinds of concepts.
On the other hand, that is an incredibly high bar.
And I'm not talking about imitation nor am I interested in semantic games, I'm talking about raw inventiveness. Not a stochastic parrot looping through a large corpus of information and a table of weights on word pairings.
Has AI ever managed to learn something humans didn't already know? It's got all the physics text books in its data set. Can it make novel inferences from that? How about in math?
Language took dozens of millennia to form, and animals have long had vocalizations. Seems like a natural building on top of existing features.
> Has AI ever managed to learn something humans didn't already know?
AlphaZero invented all new categories of strategy for games like Go, when previously we thought almost all possible tactics had been discovered. AIs are finding new kinds of proteins we never thought about, which will blow up the fields of medicine and disease in a few years once the first trials are completed.
If so, they're acting on a gigantic assumption that GPT-4 actually correctly encodes a reasonable model of the body of knowledge that went into the development of LLMs.
Help me out. Am I missing something here?
Yes the initial hypothesis that GPT-4 would know was a gigantic assumption. But a falsifiable one which we can easily generate reproducible tests for.
The idea that simulated neurons could learn anything useful at all was once a gigantic assumption too.
If you squint it's train/test separation.
But I would be very cautious about drawing conclusions from any individual neuron explanation generated in this way - even if it looks plausible by visual inspection of a few attention maps.
Based on my experience with model organisms (flies & rats, primarily), it is actually pretty amazing how analogous the techniques and goals used in this sort of research are to those we use in systems neuroscience. At a very basic level, the primary task of correlating neuron activation to a given behavior is exactly the same. However, ML researchers benefit from data being trivial to generate and entire brains being analyzable in one shot as a result, whereas in animal research elucidating the role of neurons in a single circuit costs millions of dollars and many researcher-years.
The similarities between the two are so clear that I noticed that in its Microscope tool [1], OpenAI even refers to the models they are studying as "model organisms", an anthropomorphization which I find very apt. Another article I saw a while back on HN which I thought was very cool was [2], which describes the task of identifying the role of a neuron responsible for a particular token of output. This one is especially analogous because it operates on such a small scale, much closer to what systems neuroscientists studying model organisms do.
[1] https://openai.com/research/microscope [2] https://clementneo.com/posts/2023/02/11/we-found-an-neuron
Lots of parallels to how our brains are thought to work.
On the other hand, I find it plausible that it's fundamentally impossible to assign some functionality to individual 'neurons' due to the following argument:
1. Let's assume that for a system calculating a specific function, there is a NN configuration (weights) so that at some fully connected NN layer there is a well-defined functionality for specific individual neurons - #1 represents A, #2 represents B, #3 represents C etc.
2. The exact same system outcome can be represented with infinitely many other weight combinations which effectively result in a linear transformation (i.e. every possible linear transformation) of the data vector at this layer, e.g. where #1 represents 0.1A + 0.3B + 0.6C, #2 represents 0.5B+0.5C, and #3 represents 0.4B+0.6C - in which case the functionality A (or B, or C) is not represented by any individual neurons;
3. When the system is trained, it's simply not likely that we just happen to get the best-case configuration where the theoretically separable functionality is actually separated among individual 'neurons'.
Biological minds do get this separation because each connection has a metabolic cost; but the way we train our models (both older perceptron-like layers, and modern transfomer/attention ones) do allow linking everything to everything, so the natural outcome is that functionality simply does not get cleanly split out in individual 'neurons' and each 'neuron' tends to represent some mix of multiple functionalities.
I also found this amusing. But you are loosely correct, AFAIK. GPT-4 cannot reliably explain itself in any context: say the total number of possible distinct states of GPT-4 is N; then the total number of possible distinct states of GPT-4 PLUS any context in which GPT-4 is active must be at least N + 1. So there are at least two distinct states in this scenario that GPT-4 can encounter that will necessarily appear indistinguishable to GPT-4. It doesn't matter how big the network is; it'll still encounter this limit.
And it's actually much worse than that limit because a network that's actually useful for anything has to be trained on things besides predicting itself. Notably, this is GPT-4 trying to predict GPT-2 and struggling:
> We found over 1,000 neurons with explanations that scored at least 0.8, meaning that according to GPT-4 they account for most of the neuron’s top-activating behavior. Most of these well-explained neurons are not very interesting. However, we also found many interesting neurons that GPT-4 didn't understand. We hope as explanations improve we may be able to rapidly uncover interesting qualitative understanding of model computations.
1,000 neurons out of 307,200--and even for the highest-scoring neurons, these are still partial explanations.
Reminds me of this Sam Altman quote from 2019:
"We have made a soft promise to investors that once we build this sort-of generally intelligent system, basically we will ask it to figure out a way to generate an investment return."
I don’t mean it’s not useful entirely, but I mean. It’s not useful in that it’s not deterministic enough to be trustworthy, it’s dangerous and really hard to scale therefore it’s more of an academic project than something that will make Altman as famous as Sergey Brin.
I personally take people like Hinton seriously too and think people playing with these things need more oversight themselves.
The author definitely tries to up the mysticism knob to 11 though, and the post itself is so long, you can hardly finish it before seeing this obvious critique made in the comments.
Perhaps they are all on stimulants!
The point of this post that we are commenting under is that they made this association public, at least in the neuron->token direction. I was thinking some hacker (like on hacker news) might be able to make something that can reverse it to the token->neuron direction using the public data so we could see the petertodd associated neurons. https://openaipublic.blob.core.windows.net/neuron-explainer/...
Would be interesting to try, though. I think it's likely that, due to the way glitch tokens happen, petertodd is probably an input neuron that is very randomly connected to a bunch of different hidden neurons. So it introduces some bizzare noise into a bunch of areas of the network. It's possible that some of these neurons are explainable on their own, but not within the broader context of petertodd.
Imagine telling someone in the middle of 2020, that in three years a computer will be able to speak, reason and explain everything as if it was a human, absolutely incredible!
https://www.bbc.com/news/technology-51064369
https://link.springer.com/article/10.1007/s13347-020-00396-6
https://blog.re-work.co/ai-experts-discuss-the-possibility-o...
For example, look at https://openaipublic.blob.core.windows.net/neuron-explainer/...
It's described as "expressions of completion or success" with a score of 0.38. But going through the examples, they are very consistently a sort of colloquial expression of "completion/success" with a touch of surprise and maybe challenge.
Examples are like: "Nuff said", "voila!", "Mission accomplished", "Game on!", "End of story", "enough said", "nailed it" etc.
If they expressed it as a basket of words instead of a sentence, and could come up with words that express it better I'd score it much higher.
I feel like this isn't a Yud-approved approach to AI alignment.
Ie, I think it's not that this shouldn't be done. This should certainly be done. It's just that so many more things than it should be done before we move forward.
Your comment triggered a random thought: A perfect name for Yudkowsky et al and the AGI doomers is... wait for it... the Yuddites :)
https://en.wikipedia.org/wiki/Eliezer_Yudkowsky https://twitter.com/ESYudkowsky https://www.youtube.com/watch?v=AaTRHFaaPG8 (Lex Fridman Interview)
I'm willing to bet the future of our species on my consistent victory in these types of matches, in fact.
Take this blog post for example, which between the lines reads: we don't expect to be able to align these systems ourselves, so instead we're hoping these systems are able to align each other.
Consider me not-very-soothed.
FWIW, there are plenty of AI experts who have been raising alarms as well. Hinton and Christiano, for example.
And what does charisma of AI alignment folks have to do with anything?
Nuclear weapons proliferated explicitly because they proved their scariness.
I'm not sure Yudkowski is an EA, but the EAs want him in their polycule.
I̵t̵'̵s̵ ̵o̵n̵l̵y̵ ̵b̵e̵c̵a̵u̵s̵e̵ ̵o̵f̵ ̵a̵ ̵q̵u̵i̵r̵k̵ ̵o̵f̵ ̵A̵d̵a̵m̵W̵ ̵t̵h̵a̵t̵ ̵t̵h̵i̵s̵ ̵i̵s̵ ̵p̵o̵s̵s̵i̵b̵l̵e̵ ̵a̵t̵ ̵a̵l̵l̵,̵ ̵i̵f̵ ̵G̵P̵T̵-̵2̵ ̵w̵a̵s̵ ̵t̵r̵a̵i̵n̵e̵d̵ ̵w̵i̵t̵h̵ ̵S̵G̵D̵ ̵a̵l̵m̵o̵s̵t̵ ̵n̵o̵ ̵n̵e̵u̵r̵o̵n̵s̵ ̵w̵o̵u̵l̵d̵ ̵b̵e̵ ̵i̵n̵t̵e̵r̵p̵r̵e̵t̵a̵b̵l̵e̵.̵
EDIT: This last part isn't true. I think they are only looking at the intermediate layer of the FFN which does have a privileged basis.
it does?
- The software powering the research paper
- The research itself (holy moly! They're showing the neurons!)
https://openaipublic.blob.core.windows.net/neuron-explainer/...
An interesting analogue. I think we simply aren't going to reason about the internals of neural networks to analyze safety for driving, we're just going to measure safety empirically. This will make many people very upset but it's the best we can do, and probably good enough.
I often think, “maybe I should use ChatGPT for this” then I realise I have very little way to verify what it tells me and as someone working in engineering, If I don’t understand the black box, I just can’t do it.
I’m attracted to open source, because I can look at the code understand it.
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3196841/ https://pure.uva.nl/ws/files/25987577/Split_Brain.pdf https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4204522
We can't recreate previous mental states, we just do a pretty good job (usually) of rationalizing decisions after the fact.
I think this constant degradation of humans is really foolish and harmful personally. “We’re just black boxes etc”, we might not know how brains work but we do and can understand each other.
On the other hand I’m starting to feel like “AI researchers” are the greatest black box I’ve ever seen, the logic of what they’re trying to create and their hopes for it really baffle me.
By the way, I have infinitely more hope of understanding an open source black box compared to a closed source one?
1. Don't worry, LLMs will be held accountable eventually. There's only so much embodiment and unsupervised tool control we can grant machines before personhood is in the best interests of everybody. May be forced like all the times in the past but it'll happen.
2. not every use case cares about accountability
3. accountability can be shifted. we have experience.
>I think this constant degradation of humans is really foolish and harmful personally.
Maybe you think so but there's nothing degrading about it. We are black boxes that poorly understand how said box actually works even if we like to believe otherwise. Don't know what's degrading about stating truth that's been backed by multiple studies.
Degrading is calling an achievement we hold people in high regard who accomplish stupid because a machine can do it.
>By the way, I have infinitely more hope of understanding an open source black box compared to a closed source one?
Sure i guess so.
Would you prefer a PDF?
(I'm always fascinated to hear from people who would rather read a PDF than a web-native paper like this one, especially given that web papers are actually readable on mobile devices. Do you do all of your reading on a laptop?)
A PDF is a text document that includes all the text, images, etc within it in the state you are going to perceive them. That web page is just barely even a document. None of it's contents are natively within it, it all requires executing remote code which pulls down more remote code to run just to get the actual text and images to display... which they don't in my browser. I just see an index with links that don't work and the the "Contributions" which for some reason was actually included as text.
Even as the web goes up it's own asshole in terms of recursive serial loading of javascript/json/whatever from unrelated domains and abandons all backwards compatibility, PDF, as a document, remains readable. I wish the web was still hyperlinked documents. The "application" web sucks for accessibility.
So yes, I would prefer a PDF and have a guarantee that it will look the same no matter where I read it.
Yes, I was just reading the paper and some of the javascript glitched and deleted all the contents of the document except the last section, making me lose all context and focus. Doesn't really happen with PDF files.
That feels like a loaded phrase. Is it "false confusion" adjacent?
Nothing would suggest this should work in practice, yet it just… does. In more or less zero shot. With a completely different underlying model. That’s fascinating.
Then you get a dictionary/index of LLMs
Could this be used to parallelize training?
Or create lighter overall language models?
The above would be like doing a “map”, how would we do a “reduce”?
I’m seeing they had gpt4 label every neuron but how?
We should ask AI, how are you doing this?
it would be great to see all the things theyve found for different layers
Let's see if that's the last requisite for exponential AGI growth...
Singoolaretee here we go..............
...our approach to alignment research: we want to automate the alignment research work itself. A promising aspect of this approach is that it scales with the pace of AI development. As future models become increasingly intelligent and helpful as assistants, we will find better explanations.
The distance between "better explanations" and using that as input of prompts that would automate self-improve is very small, yes?