What is a transformer model? (2022)
blogs.nvidia.com
blogs.nvidia.com
Just a few more labels, making the implicit explicit, would make it far more intelligible. Plus, last time I went through it Im pretty sure that there's either a swap on the order of the three inputs between different figures, or that it's incorrectly diagrammed.
[0]: The original paper: https://arxiv.org/abs/1706.03762
[1]: Full walkthrough for building a GPT from Scratch: https://www.youtube.com/watch?v=kCc8FmEb1nY
[2]: A simple inference only implementation in just NumPy, that's only 60 lines: https://jaykmody.com/blog/gpt-from-scratch/
[3]: Some great visualizations and high-level explanations: http://jalammar.github.io/illustrated-transformer/
[4]: An implementation that is presented side-by-side with the original paper: https://nlp.seas.harvard.edu/2018/04/03/attention.html
Bu the end of you would have done stuff like hand working out back-propagation though sums, broadcasting, batchnorm etc. Fairly intense for a regular programmer!
But even the crappy (small 500k param IIRC) Transformer model trained on a free colab in a couple of minites was relatively impressive. Looking at only 8 chars back and train on a HN thread it got the structure / layout of the page pretty good, interspersed with drunken looking HN comments.
https://youtu.be/wzfWHP6SXxY?t=4366
https://youtu.be/gKD7jPAdbpE (up to 25:42)
When I try to use transformers or any AI thing on a toy problem I come up with, it never works. And there's this blackbox of training that's hard to debug into. Yes, for the available resources, if you pick the exact problem, the exact NN architecture and exact hyperparameters, it all works out. But surely they didn't get that on the first try. So what's the tweaking process?
https://karpathy.github.io/2019/04/25/recipe/
but the general idea of "get something that can overfit first" is probably pretty good.
In my experience getting the data right is probably the most underappreciated thing. Karpathy has data as step one, but in my experience, also data representation and sampling strategy does quite the miracle.
In Part II of our book we do an end-to-end project including e.g. a moment where nothing works until we crop around "regions of interest" to balance the per-pixel classes in the training data for the UNet. This has been something I have pasted into the PyTorch forums every now and then, too.
I think I'm still at a step before the overfit. It doesn't converge to a solution on its training data (fit or overfit). And all my data is artificially generated so no cleaning is needed (though choosing a representation still matters). I don't know if that's what you mean by getting the data right or something else. Example problems that "don't work": fizzbuzz, reverse all characters in a sentence.
Vision transformers won't treat the image as a sequence of pixels but that's mostly because doing that gets very expensive very fast. The image is split into patches and the patches have positional embeddings.
Layers and lots of training. Upper layers can capture/recognize large scale features
Pixels don't get fed into transformers but that's more expense than anything. Transformers need to understand how each piece relates to every other piece. That gets very costly very fast when the "pieces" are pixels. Images are split into patches instead with positional embeddings.
as for how it learns the representations anyway. Well it's not like there's any specific intuition to it. after all, the original authors didn't anticipate the use case in Vision.
and the fact that you didn't need CNNs to extract features first didn't really come into light till this paper - https://arxiv.org/abs/2010.11929
It basically just comes down to lots of layers and training
As for > after all, the original authors didn't anticipate the use case in Vision.
To quote the original "Attention is all you need" paper: "We are excited about the future of attention-based models and plan to apply them to other tasks. We plan to extend the Transformer to problems involving input and output modalities other than text and to investigate local, restricted attention mechanisms to efficiently handle large inputs and outputs such as images, audio and video"
To say that the original authors did not anticipate the Transformers use in Vision is false.
Doing it this way is obviously throwing away a lot of information, but I assume the advantages of using a transfomer outweigh the disadvantages.
Prompted by some of the other replies I’ve read up a bit on positional embeddings. That they are used even with text transformers which otherwise lack the linearity of RNNs etc helps tremendously to clarify things.
It feels like a joke, but the fact they are mentioning GPT-3 and Megatron-Turing as the hottest new things makes this piece seem so outdated.
And as they say, attention is all you need
> researchers are studying ways to eliminate bias or toxicity if models amplify wrong or harmful language. For example, Stanford created the Center for Research on Foundation Models to explore these issues
We've seen numerous times now, censoring/railroading degrades quality. This is why SD 2 was so bad at the human form over 1.4
I get they want a PC AI but this isn't the way.
General learning + search methods that scale with data + compute win out.
Transformers in a sense is one of the most general algorithms able to tackle a whole bunch of domains from audio, video, images, language.
I’m sure another better super algorithm will be invented but Transformer is currently what we have.
During the forward pass a "query" is "created" for each token using the query projection (again, one projection per head, all the tokens run through the same projection). The keys are created using the key projection, the values are created using the dot product of query with the keys for each "token" then projected using the value projection.
But again, different models do different things. Some models bring in the positional encoding into the attention calculation rather than adding it in earlier. Practically every combination of things has been tried since that paper was published.
Models have the initial N embeddings passed through, processed by the attention+linear blocks which add something to the input. Sort of memory bus. After the last block we still have array of the same dimensions like N initial embeddings, but they mean something else now. How do we select the next token, based on the last element in that array?
Another question, that bus technically doesn't have to have the same width, right? It should be possible, for the same model's size, to trade bus width for the number of heads. Even have sort of U-Net.
And the last, on each loop resulting embedding (1) is converted into token, which is being added to the input. Converted to embedding first, it should be possible to just reuse embedding (1), and use token only for user output, right?
PS: not sure about every combination tried, we just started. Some of the problems still don't have satisfying solutions. Like hallucinations, online training. Or responsibility.
In your analogy, the bus could be any width. In practice people tend to trade "bus width" for heads, but it need not be that way.
I'm not sure understand the last part.
It's quite easy to try stuff with LLMs/transformers. The fact that a paper hasn't been written on every combination doesn't mean the haven't been tried in some way. It's not as though the architecture is the only thing.
I wish there were some way for it to communicate that certain responses about itself were more or less hardcoded.