Today's incredible state-of-the-art does not exist without the transformer architecture. Transformers aren't merely some lucky passengers riding the coattails of compute scale. If they were, then the ChatGPT app which set the world ablaze would've instead been called ChatMLP, or ChatCNN. But it's not. And in 2024 we still have no competing NLP architecture. Because the transformer is a genuinely profound, remarkable idea with remarkable properties (e.g. training parallelism). It's easy to downplay GPTs as a mostly derivative idea with the benefit of hindsight. I'm sure we'll perform the same revisionist history with state-space models, or whatever architecture eventually supplants transformers. Do GPTs build on prior work? Do other approaches and ideas deserve recognition? Yeah, obviously. Like...welcome to science. But the transformer's architects earned their praise -- including via this article -- which isn't some slight against everyone else, as if accolades were a zero-sum game. These 8 people changed our world and genuinely deserve the love!
Is there any good summary of the history of AI/deep learning from, say, late 00s/2010 to the present? I think learning some of this history would really help be better understand how we ended up at the current state of the art.
The course starts all the way back from basic statistics and goes through things like linear regression and supposedly will arrive at neural networks and machine learning at some point.
So I don't know if something like this is exactly what you're looking for, but I think that, in general, if one wants to learn about (the history) AI, then it might be a good idea to start from statistics and learn about how we got from statistics to where we are now.
No, we do. State space models are both faster and scale just as well. E.g., RWKV and Mamba.
> Transformers aren't merely some lucky passengers riding the coattails of compute scale.
Err... they are, though. They were just the type of model the right researchers were already using at the time, probably for translating between natural languages.
The bitter lesson [0] strikes again.
[0] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
Personally, I think both compute and NN architecture is probably needed to get closer to AGI.
I believe energy minimization is literal, just look at the size of that thing and imagine the power bill.
However, some advances can have huge consequences to the field compared to others, even if at the technical level they appear comparable.
One example that comes to mind is CRISPR.
And of course, a triggered Yann[2] (who is absolutely right).
But it is odd since it is actually a highly discussed topic, the history of attention. It's been discussed on HN many times before. And of course there's Lilian Weng's very famous blog post[3] that covers this in detail.
The word attention goes back well over a decade and even before Schmidhuber's usage. He has a reasonable claim but these things are always fuzzy and not exactly clear.
At least the article is more correct specifying Transformer rather than Attention, but even this is vague at best. FFormer (FFT-Transformer) was a early iteration and there were many variants. Do we call a transformer a residual attention mechanism with a residual feed forward? Can it be a convolution? There is no definitive definition but generally people mean DPMHA w/ skip layer + a processing network w/ skip layer. But this can be reflective of many architectures since every network can be decomposed into subnetworks. This even includes a 3 layer FFN (1 hidden layer).
Stories are nice, but I think it is bad to forget all the people who are contribution in less obvious ways. If a butterfly can cause a typhoon, then even a poor paper can contribute to a revolution.
[0] https://www.ft.com/content/37bb01af-ee46-4483-982f-ef3921436...
[1] https://www.bloomberg.com/opinion/features/2023-07-13/ex-goo...
[2] https://twitter.com/ylecun/status/1770471957617836138
[3] https://lilianweng.github.io/posts/2018-06-24-attention/