The Transformer Family
lilianweng.github.io
lilianweng.github.io
Can anyone recommend some great courses or other online resources for getting up to speed on the state-of-the-art with respect to AI? Not really so much looking for an "ELI5" but more of a "you have a strong programming and very-old-school AI background, here are the steps/processes you need to know to understand modern tools".
Edit: thanks for all the great replies, super helpful!
You can quickly get overwhelmed by the million good resources out there so I'll keep it to these three. If you have a strong CS background, they'll take you a long way:
(1) Transformers from Scratch: https://peterbloem.nl/blog/transformers
(2) Attention Is All You Need: https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547de...
(3) Formal Algorithms for Transformers: https://arxiv.org/abs/2207.09238
And subsequent follow-ups (ROME, editing transformer arch): https://youtu.be/_NMQyOu2HTo
I find the channel amazing explaining super complex topics in simple enough terms for people who have some background in AI.
"LSTM is dead. Long live transformers!" (Leo Dirac): https://www.youtube.com/watch?v=S27pHKBEp30
I’m also about to read the transformer chapter from this excellent upcoming book by Simon Prince:
They'll also be releasing the "From Deep Learning Foundations to Stable Diffusion" course soon, which is basically Part 2 of the course.
Even if this linked post, they have a "notations" section at the top. Almost immediately, they start using a value k that isn't defined anywhere.
https://www.manning.com/books/deep-learning-with-python-seco...
Written by François Chollet, the creator of Keras
https://github.com/rstebbing/workshop/tree/main/experiments/...
Once you understand how everything is operations on graphs, most kernels and functions become very easy to understand.
On the other hand:
> There are various forms of attention / self-attention, Transformer (Vaswani et al., 2017) relies on the scaled dot-product attention: given a query matrix , a key matrix and a value matrix , the output is a weighted sum of the value vectors, where the weight assigned to each value slot is determined by the dot-product of the query with the corresponding key
There HAS to be a better way of communicating this stuff. I'm honestly not even sure where to start decoding and explaining that paragraph.
We really need someone with the explanatory skills of https://jvns.ca/ to start helping people understand this space.
I agree that there probably could be a better "on ramp" into this material than "take an undergraduate linear algebra course", but ultimately it is a mathematical model and you're going to have to deal with the math at some point if you want to actually understand what's going on. Linear algebra and calculus are entry-level table stakes for understanding how machine learning works, and there's really no way around that.
Just like if you want to really learn how programs work you can't refuse any explanation that talks about "variables" or "functions" because that's jargon.
You can explain it, but it's going to be more at the level of "the network looks at the words" type explanation.
The analogy in my mind is this: "burning natural oil/gas completely beats figuring out cleaner & more sustainable energy sources"
My point is that "more data" here simply represents the mental effort that has already been exerted in pre-AI/DL era, which we're now capitalizing on while we can. Similar to how fossil fuels represent the energy storage efforts by earlier lifeforms that we're now capitalizing on, again while we can. It's a system way out of equilibrium, progressing while it can on borrowed resources from the prior generations.
In the long run, the AI agents will be less wasteful as they reach the limits of what data or energy is available on the margins to compete within themselves and to reach their goals. It's just we haven't reached that limit yet, and the competition at this stage is on processing more data and scaling the models at any cost.
Not really. It's not simply that modern architectures are not adding additional inductive biases, they are actively throwing away the inductive bias that used to be used by everyone. For example, it was taken for granted that you should use CNNs to give you translation invariance, but apparently now visual transformers can match that performance with the same amount of compute.
Vision transformers outperform CNNs in HUGE data regimes. On small datasets, CNNs still shine.
Also, if you take a CNN with modern tricks, they can be on par with vision transformers, e.g convnext
Transformers really dominate when you scale the amount of data to infinity
Actually, BERT is an encoder-only architecture, not decoder-only. Aside from trying to solve the same problem, GPT and BERT are quite different. This kind of confusion on now "classic" transformer models makes me kind of dubitative that the more recent and exotic ones are described very accurately...
(Clicking on the link with more details on BERT actually doesn't dispel much of the confusion; it stresses the fact that unlike GPT it's bidirectional, and indeed bidirectional is the "B" in BERT, but that's quite a disingenuous choice of terms itself - it's not "bidirectional" as in Bi-LSTM, that go left-to-right and right-to-left separately, it does the whole sequence at once; that was the real innovation of BERT).
Scrolling down to Transformer-XL starts talking about segments, from the context I _think_ it means that the input text is split into segments that are dealt with separately to cut down on the O(N^2) dependency of the transformer, but I would have assumed this kind of information to be written in a survey article.
IMHO, review articles are really great and useful, because they allow to cut through the BS that every paper has to add to get published, unify notations, and summarize the main points clearly. This article does a commendable job on the second point and, partly, on the first, but sadly lacks the third. Given the enormous task that it certainly was to compile this list, it would probably have profited from treating fewer models but putting things a bit more into perspective...
Andrej Karpathy's GPT video is a must have companion for this https://youtu.be/kCc8FmEb1nY I was going nuts trying to grok Key, Query,Position and Value until Andrej broke it down for me.