Based on my experience with model organisms (flies & rats, primarily), it is actually pretty amazing how analogous the techniques and goals used in this sort of research are to those we use in systems neuroscience. At a very basic level, the primary task of correlating neuron activation to a given behavior is exactly the same. However, ML researchers benefit from data being trivial to generate and entire brains being analyzable in one shot as a result, whereas in animal research elucidating the role of neurons in a single circuit costs millions of dollars and many researcher-years.
The similarities between the two are so clear that I noticed that in its Microscope tool [1], OpenAI even refers to the models they are studying as "model organisms", an anthropomorphization which I find very apt. Another article I saw a while back on HN which I thought was very cool was [2], which describes the task of identifying the role of a neuron responsible for a particular token of output. This one is especially analogous because it operates on such a small scale, much closer to what systems neuroscientists studying model organisms do.
[1] https://openai.com/research/microscope [2] https://clementneo.com/posts/2023/02/11/we-found-an-neuron
Lots of parallels to how our brains are thought to work.
On the other hand, I find it plausible that it's fundamentally impossible to assign some functionality to individual 'neurons' due to the following argument:
1. Let's assume that for a system calculating a specific function, there is a NN configuration (weights) so that at some fully connected NN layer there is a well-defined functionality for specific individual neurons - #1 represents A, #2 represents B, #3 represents C etc.
2. The exact same system outcome can be represented with infinitely many other weight combinations which effectively result in a linear transformation (i.e. every possible linear transformation) of the data vector at this layer, e.g. where #1 represents 0.1A + 0.3B + 0.6C, #2 represents 0.5B+0.5C, and #3 represents 0.4B+0.6C - in which case the functionality A (or B, or C) is not represented by any individual neurons;
3. When the system is trained, it's simply not likely that we just happen to get the best-case configuration where the theoretically separable functionality is actually separated among individual 'neurons'.
Biological minds do get this separation because each connection has a metabolic cost; but the way we train our models (both older perceptron-like layers, and modern transfomer/attention ones) do allow linking everything to everything, so the natural outcome is that functionality simply does not get cleanly split out in individual 'neurons' and each 'neuron' tends to represent some mix of multiple functionalities.
I also found this amusing. But you are loosely correct, AFAIK. GPT-4 cannot reliably explain itself in any context: say the total number of possible distinct states of GPT-4 is N; then the total number of possible distinct states of GPT-4 PLUS any context in which GPT-4 is active must be at least N + 1. So there are at least two distinct states in this scenario that GPT-4 can encounter that will necessarily appear indistinguishable to GPT-4. It doesn't matter how big the network is; it'll still encounter this limit.
And it's actually much worse than that limit because a network that's actually useful for anything has to be trained on things besides predicting itself. Notably, this is GPT-4 trying to predict GPT-2 and struggling:
> We found over 1,000 neurons with explanations that scored at least 0.8, meaning that according to GPT-4 they account for most of the neuron’s top-activating behavior. Most of these well-explained neurons are not very interesting. However, we also found many interesting neurons that GPT-4 didn't understand. We hope as explanations improve we may be able to rapidly uncover interesting qualitative understanding of model computations.
1,000 neurons out of 307,200--and even for the highest-scoring neurons, these are still partial explanations.