I struggle to understand why this thing works the way it does. It's possible that Vaswani et al. have made one of the greatest discoveries of this century that solved the language problem in an unintuitive, and yet very unappreciated way. It's also possible that there are other architectures that can simulate the same level of intelligence with such large numbers of parameters.
I think you re right that it's not intuitive, it's like basic arithmetic is laughing at us
The other clear benefit of transformers over an arch like RNNs (and what has probably made more of a difference imo) is that its properly parallelizable, which means you can do huge training runs in a fraction of the time. RNNs might be able to get to a level of coherence that approaches GPT-3, but with current hardware that would be very time-prohibitive.
And if I tell it something that was excatly in it's trained context windows, I get the most likely next word and the one after itm
But what happens if I ask it something slighty different than it's training context ? Or something largely different?
It's not possible for you to ask it things even slightly different from it training data, unless you ask exclusively in emojis that didn't exist yet when it was trained (in which case it sees nothing, just like when someone sends you an emoji your phone doesn't support).
Any novel sentence and even novel words like "Blobdarfnk" ARE in its training data. "Blobdarfnk" is encoded as the five tokens Bl, ob, dar, fn, and k.
In fact the projection operations are the only learned part of a Transformer's self-attention function -- the rest of self-attention is just a weighted sum of the input vectors, where the weights come from the (scaled) vector correlation matrix.
I'm not in this field but have recently found myself going on the deepest dive possible into it as my small brain can absorb.
I now know about (on a surface level) neural networks, transformers, attention mechanisms, vectors, maticies, tokenization, loss functions and all sorts of other crazy stuff.
I come out of this realizing that there are some incredibly brilliant minds behind this. I knew AI was a complex subject but not on the level I've learned about now. To get what is essentially matrix multiplications to learn complex patterns and relationships in language is mind-blowing.
And it's creative. It can have a rap battle with an alter-ego, host a quiz party with other AIs of varying personalities, co-author a short story with me, respond to me only in emojis. The list is seemingly endless. Oh, and it can also do useful things. It's my programming companion too.
And we're just getting started.
Given that Lacan already proposed the unconscious as structured language-like more than half a century ago and described attention in his turn on Freud's impulse in favor of his concept of derive, we may say, this is pretty much where our own demons live.
(I actually do think that revisiting Lacan in this context may be productive.)
(There had been times, when linguistics were still a major entry path into computing, where things were a bit different. Notably, this were also the times, which gave rise to most of the general paradigms. A certain amount of generality was even regarded a prerequisite to programming. Particularly, HN is such a great place, because it holds up this notion of generality.)
Well, if you're in need of an established theory of (semantically driven) talking machines and what derives from this, and what this may mean for us in terms of freedom, look no further.
(Mind that this is trying to talk about what's beyond/below language, necessarily using language just the same, which is – at least according to (the early) Wittgenstein – somewhat an impossibility. You can only show these things, so it takes several approaches from several directions. But there is actually something like a concise corpus of theory eventually emerging from this. Moreover, this – being transcripts of seminars – addresses an audience that is already familiar with Freud, in order to reframe this. – This is also one of the major issues with Lacan and his reception: it takes some serious investment to get into this, and this also used to have some worth on the academic markets. On the other hand, this (academic) value became indeed inflated and eventually devalued, to the point of those, who never bothered to invest, happily triumphing. Think the Great North-American Video Game Crash. But this really shouldn't be the end to what may be one of the major approaches towards what language actually means to us. The expectation that everything can be addressed directly and without prerequisites, regardless of the complexity, may actually not be met. On the other hand, there will be also never be a single "master", who is always right and without failure, bearing always the most distilled emanation of truth in their very word. – I'm also not arguing that everybody is now to become a scholar of Lacan. Rather, we may have an informed expert discussion, what may gained from this from a current perspective. E.g., if Lacan actually had something to say about an impulse-like directional vector emerging from attention (as a form of selectional focus on a semantic field), is there something to be learned from this, or, to be aware of?)
I was browsing in the medical school bookstore in Berlin (Humboldt/Charité) looking through the psychiatry section and (not joking) a third of the books were by Lacan. Will try with residual trepidation.
The reading list is already too deep and broad for this mortal. But G. Buzsaki, P. Churchkand, A. Damasio, D. Dennett, M. Donald, J. Hawkins, D. Hofstadter, C. Koch, R. Llinas, M. Minsky, Tommasi, J. Panksepp, Piaget, E. Pöppel … do find good traction for those of us who are neuroscientists interested in levels of compute that lead to language generation by human wetware.
Please end our strange fascination with fashionable nonsense. Freud was wrong. There is no Oedipus complex. Everything lacan proposed was wrong. Deleuze and Guattari's mental health clinic failed spectacularly, and Deleuze ended up killing himself at the end (supposedly due to back pain?)
They literally describe their thought as being "Schizoanalysis". How many more red flags do you need?
Also, the more "modern" takes on this from techno folks, such as from Nick Land (Fanged Noumena), are openly fascist - https://en.wikipedia.org/wiki/Dark_Enlightenment
If you want cultural critique from smart people without it turning into fashionable nonsense, I recommend Mark Fischer, but be warned, he too killed himself.
Regarding charlatans, mind that there are already few who have actually studied this. (I'm one of them.)
Regarding Lacan, he provides us with an established theory of "talking machines", and, in a philosophical context, how they relate to our very freedom (or, what freedom may even be). This isn't totally useless in our current situation, and NB, it's actually quite the opposite of fascism.
That's just to correct the record. I have no desire to re-litigate Sokal/Bogdanoff and so on. Good day sir cheerio.
(A turn towards the dogmatic is something I'm pretty much expecting from the current launch of AI anyway, simply, because the productions systematically favor the semantic center. So it may be worth putting some generality against this, rather than being overly selective.)
It's a bit more plausible when we phrase it that way...
I agree with the sentiment that each individual dimension isn't meaningful, and I also feel like it's misleading for the article to frame it that way. But there's a grain of truth: the last step to predicting the output token is to take the dot product between some embedding and all the possible tokens' embeddings (we can interpret the last layer as just a table of token embeddings). Taking dot products in this space are equivalent to comparing the "distance" between the model's proposal and each possible output token. In that space, words like "apple" and "banana" are closer together than they are to "rotisserie chicken," so there is some coarse structure there.
Doing this, we gave the space meaning by the fact that cosine similarity is meaningful proxy for semantic similarity. Individual dimensions aren't meaningful, but distance in this space is.
A stronger article would attempt to replicate the word2vec analogy experiments (imo one of the more fascinating parts of that paper) with GPT's embeddings. I'd love to see if that property holds.