And how does the NN represent the token at the output layer? Is it a binary representation of the token number?
Or does it have a neuron for each token it knows and ChatGPT takes the most activated neuron as the answer?
But I personally think it's a coincidence, and it just so happens that 50k tokens are enough for the level of complexity the models have right now.
Tokens are part of words, approx 4 characters or 75% of word.
It gives a list of tokens with their probabilities on output.
It's a short list with highest probabilities.
Temperature controls which tokens to pick - usually 0% = top one only (consistent results), closer to 100% means more randomness (more "creativity").
It's "glorified markov chain" in the same sense that sqlite is just "glorified bubble sort".
For example: if you limited all inout/output to the same 100 words, could you stay within the token limit permanently?