Cynthia Rudin and interpretable ML models
quantamagazine.org
quantamagazine.org
At a very low level, there’s no secret: all of the weights are there, all of the relationships between weights are known. The problem is that it doesn’t tell you anything about the emergent properties of the network, in the same way that quantum physics doesn’t give much insight into biology.
It may be possible that there is no English sentence you can utter about the network which is both explanatory and fully accurate. What is the network doing? It’s trying to approximate the function you’ve given it. That’s it.
You can try other things like ablation to find the effects of lobotomizing the network in certain ways, but this also can’t fully explain 2nd and higher order relationships.
To my knowledge, some angles have been explored. Finding which inputs maximally stimulate indiv neurons, and tracing that up the chain.
There was a fantastic article that did this on ImageNet - the takeaway was that neurons adjacent to the input encode sharp, basic features (like edges), and "higher-level" ones encode more nuanced stuff (textures, fur, wheels).
Then ablations, as you mentioned. And finally, rerunning the same training data on differently sized + shaped architectures.
AI systems are missing at least one level of complexity. All neurons in a NN layer fire simultaneously, unlike brain neurons, which trigger each other async.
Afaik the properties you’re referring to on AlexNet (the network trained to perform on Image Net) have to do with the nature of repeated convolution operations which, while interesting, is not as deeply insightful as I believe the article is aiming for.
FTFY. Philosophy aside, humans understand complex things by building simpler abstractions for them. I think the key to NN understandability centers on building simpler NNs focused on specific tasks that build into bigger ones.
> The explanations have to be wrong, because if their explanations were always right, you could just replace the black box with the explanations.
I've never thought of it like that, but I think that really gets to the core of the issue...
To be able to reason about the behaviour of a neural network with 100% confidence you would need an explanation that incorporated every weight, otherwise you have to accept some degree of uncertainty.
But the idea you could ever explain something with billions of parameters in a way that a human could comprehend seems ridiculous. Imperfect generalisation therefore would inevitably be required – "this bit of the network seems to perform x function". But in doing this you have to accept that this generalisation comes with uncertainty because such a high-level approximation is unlikely to capture all of the nuances. And if it did capture all nuances then the black box isn't needed.
Or, at more length and in more depth, The Analytical Language of John Wilkins by Borges.
> Lewis Carroll, in Sylvie and Bruno Concluded (1893), made the point humorously with his description of a fictional map that had "the scale of a mile to the mile". A character notes some practical difficulties with such a map and states that "we now use the country itself, as its own map, and I assure you it does nearly as well."
Neural networks aren’t like that. The whole thing works together simultaneously to derive a result. So I’m skeptical that there is a high level explanation other than “it’s doing it’s best to approximate the data it was trained on” which is not that useful.
Suppose you have the spinning color wheel on your MacBook. You give it to a computer scientist who says that the program is in an infinite loop, and the wheel will never stop spinning. Then you give the computer to a superhuman physicist who studies all of the components including the state of the transistors and internal components and concludes that the wheel will spin for 9 years before the screen gives out. [1]
The point is that both are correct, but we’d really prefer the first explanation, since it gives insight into what the program is doing.
The problem is that neural networks aren’t like programs. They don’t have reducible internal components unless you reduce down to the weights of the network itself. [2]
But “this portion of the network appears to recognize faces” is still a useful abstraction. In the same way it is for neurology. Even if both NNs and brains sometimes see faces that aren’t there because the network isn’t “recognizing faces” but a system which signals on a complex set of heuristics that mostly detects faces.
How does the answer change if I only want 99% confidence and accept the existence of illusions? — can we describe the bulk of the behavior in a summary?
https://www.alignmentforum.org/posts/j84JhErNezMxyK4dH/llm-m...
However, to think we cannot comprehend them at all is too much. We can certainly ”query” a NN to understand the weights and activations that led to some decision in some specific scenario. How to consume and use that information is another matter, of course, but a debugging process is certainly feasible (as it is with any large scale computer system).
Plus, the architectures of neural networks are not as complicated as most people think. Sure it gets weird after training, but the basic architecture of components are not nearly as complicated as the hardware of a GPU for example, or even that of a complex OS kernel.
reminds me of: "The Tao that can be told is not the eternal Tao."
The real problem IMO is our ability to represent natural concepts as such graphs. Is love, for example, a DFA? Can we search for its isomorphism?
This is very insightful. I've long maintained that "understanding" requires some predictive ability in addition to an ability to explain what's happened in the past.
In the case of neural networks, we should be able to "simply" generate the network that has the desired behavior.
But we can't. We're still shackled to compiling some collection of training data that may or may not produce the desired behavior, tested by some other collection of testing data that may or may not test the desired phenomenon.
Also, remember that the explanation doesn't really have to fully contain the semantics of the model. It doesn't have to fully encode the weights, just activate sufficiently sympathetic structures in a human brain.
Ask GPT4 to do a task, and then ask it to the same task showing its work; you'll find that GPT4 is less likely to make mistakes on the latter. This is especially apparent for tasks like counting # of words and multi-step problems, which GPT normally has trouble with.
But GPT4 still tends to struggle even breaking the task down, to the point where it starts producing extremely obvious mistakes (e.g. "the turtle moves 1 unit up, from (1, 0) to (2, 0)"). One possibility is that it isn't actually showing its work, it's just generating backwards explanations from a latent conclusion. Maybe this research will clarify whether this is the case, and help us develop a more coherent LLM.
I disagree with the premise here. You don't understand how all the computations work in the brains of people's predictions whom you trust. You simply have a mental calculation of their batting average through exposure to their track record and this batting average functions as a proxy for trust.
I find this is more or less the same way that I learn whether or not I can rely on GPT-4 for a particular use-case. If it's batting average is north of a certain % for a given use-case, then it doesn't need to be right 100% of the time for me to derive value from relying on it.
I think we are slowly crossing a threshold where we accept indeterminism and mistakes from machines in a way that we haven't in the past.
Otherwise, you are just trusting the training set and the training process. Sometimes that's fine, as the article also mentions.
In short: we trust people because we can model them effectively, and being one helps. We trust mathematical models and traditional programs because we understand them and can check their work in any specific instance. Large ANNs don't (yet) benefit from either of these two kinds of trust.
That said, someone has to have a good understanding of what is happening (not necessarily the end user). In my experience neural networks today are basically like compiled software which source code has been lost. Every act of debugging and fixing has to include also reverse engineering (I’m talking about the actual network weights, not the e.g. PyTorch code). It’s too inefficient and cumbersome, and as they become more and more common, statistics says the failures will become more often and more severe.
When do we, for example, decide to put the first NN on an airplane cockpit?
Okay. Since apparently no one has the explanations desired, we have to guess. So, let's do some guessing:
Given so many parameters, etc., we have in some sense -- in some case of geometry, spaces, maybe vector spaces, maybe as in linear algebra -- a lot of dimensions.
Then something surprising holds (once we get precise about a space, easy enough to prove): Given a sphere in the space, we can calculate its volume. We can do this for the space of any finite dimension. Here is the surprise: As we have a lot of dimensions, there is a LOT of volume in that sphere, and nearly all that volume is just inside the surface of that sphere. E.g., if do some work in nearest neighbors, discover this surprise in strong terms.
Net, in the space being considered, there is a LOT of volume. Then ...: There is plenty of volume to put faces of cats over here, dogs over there, men another place, women still well separated, essays on bone cancer far away, ..., for thousands, millions, ..., more things, thoughts, topics, etc. Then given some new data, say, a white cat not in the training data, likely the data on that white cat will settle on the volume with the cats instead of dogs, monkeys, etc. and, thus, we will have recognized a cat via some emergent functionality.
Just a guess.
Because we understand these functions and tables, we understand exactly how well the network will work, and also what is missing (i.e., how we can expand its accuracy.)
I think this is a very hard problem, but it is one that needs to be solved.
If you take an information-theoretic approach to it, and think of a DL model like any other model, there is a certain equivalence in understanding the model features and how it behaves with reference to the universe of data it is applied to.
It was an interesting article but I felt like it created problems that need not be there (or maybe it's just describing problems that others created?)
I think that the bulk of ML has so far produced what our brains in fact easily see. We can easily perform classification, or generation.
We’re just “meat bags”.
I believe we've already mapped some invertebrate brain, maybe 250 neurons.
Low level descriptions provide shockingly little insight into high level behavior.
Also, man, quanta feels really rough and popsci-y when it comes to CS.
I'm guessing this is just the Gell-Mann Amnesia effect. Why do you think the quality is better for other fields?