Decomposing language models into understandable components
anthropic.com
anthropic.com
Well, sort of ..., I'm refining an algo that takes several (carefully calibrated) outputs from a given LLM and infers the most plausible set of parameters behind it. I was expecting to find clusters of parameters very much alike to what they observe.
I informally call this problem inverting an LLM, and obv., it turns out to be non-trivial to solve. Not completely impossible, tho! as so far I've found some good approximations to it.
Anyway, quite an interesting read, def. will keep an eye on what they publish in the future.
Also, from the linked manuscript at the end,
>Another hypothesis is that some features are actually higher-dimensional feature manifolds which dictionary learning is approximating.
Well, you have something that behaves like a continuous, smooth space so you could define as many manifolds as you'd need to suit your needs, so yes :^). But, pedantry off, I get the idea and IMO that's definitely what's going on and the right framework to approach this problem from.
One amazing realization one can get from this is, what is the conceptual equivalent of the transition functions that connect all different manifolds in this LLM space? When you see it your mind will be blown, not because of its complexity, but rather because of its exceptional simplicity.
Could you elaborate on what you mean by "transition functions" here?
although the quoted sentence does not make sense to me; transition maps connect different patches of one manifold. it's possible the "LLM space" gp is talking about is a parameter space of some nature each of whose points is a manifold, but that seems like a stretch
Instead of "all different manifolds" I should have written "all different ways to define a manifold". Also, now that I think about it, it may not necessarily be all of them.
The important thing here is that, however you define them and their transition maps, you'll find they're very much alike. As if there was some sort of general structure that is highly preferred over others ...
what is the manifold here and what evidence do you have for this / what does it look like when you "define" them?
They also reference, https://distill.pub/2020/circuits/equivariance/.
I'd also add, https://distill.pub/2021/multimodal-neurons/.
I’m curious to learn more about how LLMs work too.
You may need some intermediate knowledge of linear algebra and this thing called "data science" nowadays, which is pretty much knowing how to mangle data and visualize it.
Try creating a small model on your own, it doesn't have to be super fancy just make sure it does something you want it to do. And then ... you'll probably could go on your own then.
But if this technique scales up, then Anthropic has fixed that. They can figure out what different groups of neurons are actually doing, and use that to control the LLM's behavior. That could help with preventing accidentally misaligned AIs.
Hm. I wish they'd said more about that. Does that mean they found the same feature recognizers when training with the same training set? Or what? This tells us something, but what does it tell us?
Typically, you can take a pre-trained model and retrain it on your new dataset by only changing the weights of the last layer(s).
Some loss functions even measures the difference between the high-level features of two images, typically extracted from a pre-trained CNN (Perceptual Loss).
[1]Matt Zeiler did an amazing work on these findings 10 years ago (https://arxiv.org/abs/1311.2901).
https://srush.github.io/raspy/
I don't know if you can integrate them into a model. I think you might run out of space, since these aren't polysemantic and so would take up a lot more "room" than learned neurons.
Edit - tokenising is a form of this, you're pre-transforming the data to save it having to learn patterns you know are important.
Because if you can see what each part is doing, then theoretically you can find ways to create just the set of features you want. Or maybe tune features that have redundant capacity or something.
Maybe by studying the features they will get to the point where the knowledge can be distilled into something more like a very rich and finely defined knowledge graph.
After all us plebs are allow to use computers and latest CPUs and internet and stuff already! Yes there is shit happening like scams, and worse but it is better than limiting what people can do.
That LLMs are capable of what they are at the compute density they are strongly signals to me that the task of making a productive knowledge worker is in overhang territory.
The missing piece isn’t LLM advancement, it’s LLM management.
Building trust in an inwardly-adversarial LLM org chart that reports to you.
We don't re-evaluate our astrophysics models when reading a cooking book.
All these LLMs appear to be converging around these features.
This research (and it’s parent and sibling papers, from the LW article) seem to be about picking out those colored graph components from the floating point soup?
edit: ah, looked at the paper, they did it unsupervised, with a sparse autoencoder.
This is even less surprising given LLMs are applied to models with a known hierarchical structure and symmetry.
Can anyone say exactly what's novel in these findings? From a layman's point of view, this sounds like announcing the invention of gunpowder.
"In physics, wherever there is a linear system with a "superposition principle", a convolution operation makes an appearance."
I'm working this out in more details but it is uncanny how much it works out.
I have a discord if you want to discuss this further
I don't see anything wildly different now, other than scale and youth and the hubris that accompanies those things.
What's your best examples of this? Some of the most impressive examples I've seen ended up being likely in the dataset, or very close to being so. I've yet to see something where it definitely wasn't approximately in the dataset and was solved in a way that seemed to use some sort of novel process, but open to being wrong.
Every single day, I get immense use out of modern language models. Even if an output is similar to something it's already processed, that's fine! Such is the nature of synthesis.
The human brain is a salient point because often we are using AI so that the human brain can do less. Get this GPU to RTFM instead of the human. The human time is more valuable. All the while making the human brain probably less effective (compare someone who learns another language vs. someone who speaks it through an AI translator only).
I hold both points of view that AI is both marvelous, but also concerning in terms of energy use.
To nitpick - in " NNs are to emergent behavior what crypto is to cash " applies more to large language models. Simpler NNs for easy tasks that don't consume much power wouldn't apply (that might be like a VISA card?)
Is it that this behavior is the result of any system at scale? That is undeniably preposterous.
Is it that the human brain is more efficient? At energy usage, sure, but for me to find an individual who is capable enough to assist me in the manner GPT does, at the speed and level of breadth and depth that it does, would be next to impossible. If I did, their required compensation would be astronomical.
What are you arguing for or against? Are you aware that these systems will, like all previous computationally intensive systems, become drastically more efficient over time?
They are not novel if there is an equivalent pattern in the training dataset. I guess you are not really trying anything that isn't available already on github or google in some form. If you think you do then please show an example of "entirely novel and original idea", that GPT-4 developed for you. I had at least 4 cases in which ChatGPT failed to produce correct solution (after pushing it for hours to correct itself in many ways) in an actual novel problem (solution not longer than 200 lines of code) for which there was no solution on Google or github. But you can't blame statistical model that was trained to create the most probable outcomes based on it's training data.
Usually, drawing from existing knowledge is the whole appeal of using GPT. Over time, you actually begin to get a sense of what the model is good at and bad at, and it's good at an incredible amount of things. I get it to write novel code constantly and I think that playing with it and confirming that for yourself is better than me showing you.
Apparently, sit there and write plain English at a billion-dollar cluster and wait for astonishing answers at 300bps.
Like, a whole decade of ML: better optimisers, better init, residual connections, tokenisation and token embeddings, training with large batches over thousands of machines, the attention mechanism, causal masking, flash attention and other memory optimisations, and even having the foresight to train on the totality of web text.
Not seeing the intermediate steps doesn't mean they are not essential and needed.
If you still believe a toy 10x10 fully connected net is the same with current models (bar scaling), then what is you opinion on MLP-Mixer? That was an "MLP is all you need" moment but it didn't lead to adoption.
Reminds me of the old dudes in the gym who come to you to tell you how they used to bench four plates when they were young. In their mind, they are badasses. In their mind only.
BTW, awesome "flex" about how much time you spend at the gym, you sound like a guy who knows what he's talking about ;)
You too might be a Hubris News celebrity one day.