How does this relate to vectors? It was my understanding that the tokens were vectors and this seems to show them as an integer.
It's probably a really obvious question to anyone who knows AI but I figured if I have it someone else does too.
How does this relate to vectors? It was my understanding that the tokens were vectors and this seems to show them as an integer.
It's probably a really obvious question to anyone who knows AI but I figured if I have it someone else does too.
More completely, you can think of the integers as being implicitly a one-hot vector encoding. So say you have a vocab size of 20,000 and you want Token #3. The one-hot vector would be a 20,000 length vector of zeros with a one in position 3. This vector is then multiplied against the embedding table/matrix. Although in practice this is equivalent to just selecting one row directly, so it's implemented as such and there's no reason to explicitly make the large one-hot vectors.
Is there some reference table somewhere mapping more code idioms like this to equivalent nn representations?
One other example would be how multi-head attention is implemented with a single matrix. You don’t actually create matrices for each of the N ‘heads’ separately. It’s a logical distinction
Other idioms I can think of, in my words:
Softmax = take the maximum (but in a differentiable way)
tanh/sigmoid/relu = a switch. "activation"
cross entropy loss = average(-log(probability you gave to the right answer)). Averaged over the current batch you are training on for this step. (Sorry that is still quite mathy).
This step is deciding which clusters of letters (or whatever) get a vector and then giving them a scalar unique ID for conveniences' sake.
The training then determines what that vector actually is.
In the case of more advanced language models like LLMs, a given token can be paired with many other features of the token (such as dependencies or parts-of-speech) to make an integer represent one of many permutations on the same word based on its usage.
And vocabulary is just an array / vector / list - it depends which programming language you use, each has each own terminology for that data structure.
For example LLaMA vocabulary has 32,000 tokens.