The GPT Architecture, on a Napkin
dugas.ch
dugas.ch
But reality is usually much messier. The real training code will be littered with hundreds of ugly little tricks to make it work. A large part of it will be input preprocessing and data engineering, tricks to deal with exploding/vanishing gradients, monitoring, learning rate schedules and optimizer cycling, complexity for distributed training, regularization tricks, changing parts of the architecture for performance reasons (like attention), and so on.
There are many rules of thumb that took the last 5+ years to discover but are now quite standard. You are nit picking on fully connected, but if we add dropout, weight initialization, and adaptive learning rate to what they said, then we are fairly close to being able at least get a deep architecture to overfit a toy dataset and be off to the races for then applying it to a larger dataset.
See this logbook from training the GPT-3 sized OPT model - https://github.com/facebookresearch/metaseq/blob/main/projec...
In the original dataset you already combine dozens of sources - web scrapes, book collections, paper collections, materials in many languages, etc. In the second stage you have thousands of small supervised datasets. In the third stage you have to label. So I think the dataset building phase is pretty difficult.
Of course, there's all the secret sauce to actually getting the models to learn anything, and all the empirical progress we make to make the training more efficient (ReLUs, etc). But how many of those are fundamental, vs. simply efficiency shortcuts? And: if you'd asked me 10 years ago what I thought it would take to get the kind of output these large models are getting these days, I would not have guessed anything nearly as simple as what those models actually are.
1 - To encode position somehow (which author details)
2 - Because sine is easy “noise” for the network to learn.
There’s also a bunch of cool tricks here even down to the PyTorch implementation to optimize this encoding by exploiting the nature of sine/cosine which is an added reason for its popularity in Transformer architectures. If you like math I recommend diving into it as it’s quick but fun!
(Side note it’s also falling out of fashion for other encoding methods. e.g. rotary positional encoding is vastly popular in the reformer branch of transformers)
My interpretation is the sinusoidal converts the values from 0..N to 0..1. This may result in more predictable changes than using large integer numbers, such that the same word in different positions doesn't lose all its meaning.
Using integers and having a layer to compute this operation could also work, so maybe this is an optimization, eg it reduces training time or yields better results.
For an example, if your training set consists of the words rain and thunder used interchangeably a lot, but the word "today" is only used once in the sentence "there is no rain today", then a Markov chain based on the data would never output "there is no thunder today", but a transformer might.
In other words, information compression (eg. equating rain with thunder) isn't just for practicability, it's a necessary requirement for (the current generation of) good language models.
The final diagram [1] is also much less "readable" than the original figures from the paper.
[1] https://dugas.ch/artificial_curiosity/img/GPT_architecture/f...
For example, how does the explanation in the article produce a module that can solve this:
“ To calculate the hypotenuse of a triangle with one side that is 12 inches long and another i S side that is 36 centimeters long, a 6th grader might say something like this: "First, we need to convert the 36 centimeters into inches so that both sides of the triangle are in the same units. We can do this by dividing 36 by 2.54, which is the number of centimeters in one inch. This gives us 14.173228 inches. Then, we can use the Pythagorean theorem to find the length of the hypotenuse. The Pythagorean theorem says that in a right triangle, the square of the length of the hypotenuse is equal to the sum of the squares of the lengths of the other two sides. So we can use this formula to find the length of the hypotenuse: a^2 + b^2 = c^2. In our triangle, the length of one side is 12 inches, and the length of the other side is 14.173228 inches. So we can plug those numbers into the formula like this: 12^2 + 14.173228^2 = c^2. Then we just need to do the math to find the value of c. 12^2 is 144, and 14.173228^2 is 201.837296. So if we add those two numbers together, we get 346.837296. And if we take the square root of that number, we get the length of the hypotenuse, which is 18.816199 inches.”
You gotta be REAL careful with ChatGPT output that sounds convincing and technical. It's very good at convincingly making stuff up, even math-y science-y sounding stuff.
Agree it's quite convincing.
In many cases each word has a number/token associated with it, but some words get broken up into several tokens. Same with longer numbers.
You can check out their tokenizer under the following domain. It probably gives you an idea why GPT is not very good at dealing with numbers.
Example 1:
Input: 6984654984165498 x 83749872394871982798
Output: The result of the multiplication is 58307472585676078521357388975872646506.
Expected: 584963963646067047073619488641103404
Example 2:
Input: 45649898465897645132987645146987156456978 x 879846498164168468465465169816989441
Output: I'm sorry, but I'm not able to perform calculations involving such large numbers. My knowledge cutoff is 2021, and I am not currently able to browse the internet, so I don't have access to any updated information. Is there something else I can help you with?
Expected: 40164903306769889413456328247690955984247293637932739659324105787247996769298
Q (Query) is like a search query. K (Key) is like a set of tags or attributes of each word that the query can look for.
Imagine the self-attention scores for the sentence ”The chicken crossed the road because it wanted to get to the other side.”
Let’s look specifically at the word “it”. We can imagine the matrix Q_it as a representation of attributes of the word we want “it” to gain context from. For example, “it” is a pronoun, and therefore gains context from nouns. Therefore, the matrix Q_it might have some representation of “noun” in it, because we want it to take context from nouns.
In other words, one of the goals during the training process of a transformer network is to train a weight W_q to map a word to a matrix representation of attributes of words it gains context from. So W_q should map any word that’s a pronoun to a noun query.
Similarly, K can be thought of as a representation of attributes of each word. So K_chicken should also have a “noun” tag. So one of the goals during training is to train W_k, that maps the latent-space representation of a word (chicken) to a matrix representation (K_chicken) of it’s important attributes (noun).
When we take the dot product of Q and K, what we’re finding is the similarity between those 2 matrices. In other words Q_it * K_chicken should have a high value because “it” is looking for nouns in it’s query, and “chicken” is a noun.
Obviously this is a very human-centric explanation and how the weights W_q and W_k are trained in practice may not align perfectly with human interpretable concepts, but hopefully helps with understanding.
Well, since its described in a paper and there are open source GPT implementations (e.g., GPT-NeoX), probably not much.