Mapping the semantic void: Strange goings-on in GPT embedding spaces
lesswrong.com
lesswrong.com
My interpretation of the finding is that the LLM training has revealed the primordial nature of human language, akin to the primary function of genetic code predominately being to ensure the ongoing mechanical homeostasis of the organism.
Language’s evolutionary purpose is to drive group-fitness. Group management fundamentals are thus the predominant themes in the undefined latent space.
Who is in our tribe? Who should we trust? Who should we kill? Who leads us?
These matters likely dominated human speech for a thousand millennia.
We've truly reached peak hype. It's a token predictor dude.
Notwrong!
Since there are ~5000 dimensions, in which of those dimensions are we moving k out ?
Is the idea you just move, out, in all dimensions such that the final Euclidean distance is k ?
Seems that’s how they get multiple samples at those distances.
Either way I think it’s more interesting to go out in specific dimensions. Ideally there is a mapping between each dimension and something inherent about the token, like the part where a dimension corresponds with the first word of the token.
We went through this discovery phase when we were generating images using autoencoders, same idea, some of those dimensions would correspond to certain features of the image, so moving along them would change the image output in some predictable way.
Either way, I think the overall structure of those spaces says something about how the human brain works ( given we invented the language). I’m interested to see if anything neurologic can be derived from those vector embeddings.
The complexity and interdependence of dimensions within embeddings make it practically impossible to ascribe specific, human-understandable meanings to individual elements or dimensions.
This research actually adds to that complexity. Rather than making the meaning of individual dimensions more understandable, it shows that the embedding space has a strange, layered structure.
It suggests a peculiar, almost nonsensical organization of concepts at different distances from the central point of typical token embeddings, which doesn't make it easier to pinpoint what each dimension means.
It emphasizes that embeddings capture information in a distributed and highly contextual manner. The fact that embeddings for "nokens" can lead to arbitrary or bizarre categorizations when taken out of the typical token zone underscores that embedding dimensions don't have straightforward, easily interpretable meanings.
Noken-space is all the undefined territory full of noise equivalent to the actual visual noise we see in Stable Diffusion when you travel along a vector away from an island of coherence. To put it poetically, concept space is a vast hyperspace sea of random garbage (like unallocated RAM) and there are tiny islands or planets of meaning and value that look like things that matter to humanity.
I don’t see the categorizations as being very bizarre. Putting on my amateur anthropology hat, each one described in the paper was clearly an important topic to a “primitive” person. Concerned with survival, people would talk about group dynamics, sharp things, plants/animals, things that look like infections (small flat round yellow white), and places. These are pretty much all you need to talk about.
Looking at an English LLM vector space to understand fundamental principles of language is like looking at human DNA and trying to understand the first eukaryotes. It’s been complicated by effectively infinite generations of specializations adding noise.
The probing work described here identifies some common principles that survive through the generations. A commenter above wondered if this was discovering something about the nature of the brain. I think it’s the nature of culture. Etymology is the product of culture and history.
> Is the idea you just move, out, in all dimensions such that the final Euclidean distance is k ?
If I understand correctly, the basic idea is that in earlier experiments, they were able to use a relatively-simple technique to come up with what they call a "probe vector" in the embedding space which represents "the property of a word starting with <letter>". For any given token, the authors established that it was 98% probable that the token's embedding vector would be closer (by cosine similarity) to the "probe vector" representing the first letter of that word than any other probe vector. This is shown in the first graph of the section "A puzzling discovery".
With that in mind, the diagram below that should start making more sense: "emb" is a particular token's embedding vector and "probe" is the probe vector for the token's first letter. "emb_proj" is the projection of "emb" onto "probe".
What they're doing is tweaking the network weights by subtracting multiples of `emb_proj` from `emb` (where the specific multiple is the parameter K), and then seeing how it behaves differently for different values of K.
Their original observation when doing this was that it reliably caused the model to claim that the first letter of the tweaked word was not the letter in question. In this article, they're trying to figure out how far they can push the tweak and still get reasonably accurate definitions of a token.
What they discovered is that when they push a token's embedding vector further and further out along its "first-letter vector" and ask the network to define that word, the definitions it provides seem to follow particular themes during different regimes of K.
I suspect they will be the “antiquated” or “child-like” concepts identified here: the things that matter to hunter-gatherer people.
It's easy to verify this for yourself in a simple example in something like python. Draw from a multivariate normal distribution (or almost anything) in a 1000-dimensional space and look at the norm of your vector. If you do this a lot of times you'll see that the answer is close to the same each time. This even holds true if you replace the vector norm by the distance to an arbitrary fixed point.
edit: here's some code:
import numpy import matplotlib.pyplot norms = [] dists = [] for i in range(5000): vec = numpy.random.normal(size=4096) norms.append(numpy.linalg.norm(vec)) dists.append(numpy.linalg.norm(vec-numpy.ones(4096))) matplotlib.pyplot.hist(norms) matplotlib.pyplot.show() matplotlib.pyplot.hist(dists) matplotlib.pyplot.show()
(didn't render right, sorry)
For instance, consider a unit cube in n dimensions. The volume of this cube (1 unit on each side) remains constant, but the distance from the center to a corner increases with the square root of the number of dimensions. This implies that in high dimensions, most of the volume of a hypercube is located at its corners. This is bizarre to comprehend because it's hard to stop imagining a "cube" in the 3D sense.
Algorithms that rely on distance measures (like k-nearest neighbors or clustering) may perform poorly in high-dimensional spaces without dimensionality reduction techniques like PCA (Principal Component Analysis) or t-SNE (t-Distributed Stochastic Neighbor Embedding). And even with these techniques, you can get very misleading results.
This is incorrect. If you put a little box of small side length x at each corner, then each box has volume x^n and there are 2^n many of them, so the total volume they take up is 2^nx^n=(2x)^n, which is very small. Maybe what you meant is that most of the volume of a high-dimensional hypersphere is near its surface.
Surface area is 2-volume.
If not, then your point isn't very clear to me.
"The phenomena discussed in the text are more about the inherent characteristics of the model's representation space rather than the input data being flawed."
It took three attempts, the first two were merely a rephrasing/paraphrasing of the abstract.
They say there are nokens which are vectors that seem to represent something because they are organized in that space in a regular manner.
They think those vectors may have something interesting about them because they seem to be somewhat organized and related to each other.
so they make new vectors
It's an article observing how GPT-J token embeddings are positioned in 'embedding space', with some connections drawn to GPT 3.
They then experiment with having GPT-J provide definitions for "nokens" (A "noken" is basically a made up embedding created by modifying the embedding generated from a real token) to see what happens when a model is presented with a novel embedding outside of the trained embedding space to see how it interprets them.
Diving in a little more:
An observation from the article is that the first letter of a word represented by an embedding is actually encoded into that embedding such that you could identify what letter any embedding starts with, with 98% certainty.
They observe this linear relationship and devise an experiment. They ask GPT-J what letter the word "icon" starts with, and it correctly replies "I". They then create a "noken" for the word "icon" and modify it so that it no longer represents that the word starts with "I" and ask GPT-J again what letter this "noken" starts with. GPT-J then incorrectly replies that the word "icon" does not start with the letter "I".
So then they pose the question "can the model still define the word 'broccoli' even if we shift the first letter away from 'B'?" and the answer is almost always yes, changing the semantic "first letter" of an embedding doesn't change the model's ability to understand the embedding.
They then do a series of experiments asking GPT-J to provide the "typical definition for the word '<noken>'." with the example word 'hate' being used. They gradually shifted the first letter away from 'H', and see that up until very high levels of modification, the model can still give a normal definition for the modified 'hate' noken ("a strong feeling of dislike or hostility.").
At extreme modifications to the noken, the model starts providing strange definitions ("a person who is not a member of a particular group.", "a period of time during which a person or thing is in a state of being").
They repeat the experiment and note that this behavior of the definition keeping stable up until 'collapse' is common across many tokens:
> ...usually ending up with something about a person who isn't a member of a group by k = 100, having passed through one or more other themes involving things like Royal families, places of refuge, small round holes and yellowish-white things, to name a few of the baffling tropes that began to appear regularly.
That said I am a noob about LLMs.
In this model (one of many infinitely many possible models) a word is a point in 4096-space.
This article is trying to tease out the structure of those points, and suggesting we might be able to conclude something about natural language from that structure - like looking at the large-scale structure of the Universe and deriving information about the Big Bang.
Obvious questions: is the large-scale structure conserved across languages?
What happens if we train on random tokens - a corpus of noise? Does structure still emerge?
It might be interesting, it might be an artifact. I'd be curious to know what happens when you only examine complete words.