Instead of coming up with more and more heuristics to chop a sequence of bytes up in "words" in a vocabulary, we could simply set a limit on the size of the vocabulary (number of tokens), put all bytes in there (so we can at least handle any input byte by byte), and pack the remaining space with the most common multi-byte byte sequences. Then you end up with tokens like here.
Generation is ultimately deterministic (seeded prng) so backtracking wouldn't make sense.
I found Stephen Wolframs explanation helpful. He has a YouTube video version which I enjoyed too. This blog post was on HN last month, but I never get good search results on hn
https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Next token prediction produces the most head exploding emergent effects.
You can put people in an fMRI and ask them to think "car".
You can ask someone to think of objects and detect when they think "car".
What happened there pairing a bunch of tensors to meanings and matching them.
We can do something similar with embeddings.
To be clear I don't intend to give the impression that these LLMs are doing something miraculous. Just that we are increasingly peeling back the veil of how brains think.
I don't know about other people, but when I think “car” really hard, I can feel the muscles in my throat adjust slightly to match the sound of the word “car”. Perhaps that sort of thing is what the MRI machines is picking up, rather than being able to pick up some kind of "internal representation" of car.
It'll also light up in the parts of my brain to do with reading, writing, hearing the word in the languages I speak.
What does car mean to me if it doesn't connect to all the concepts that relate to cars?
There's no "better way" to do it because the tokens are all meaningless to ChatGPT, it only cares about how efficiently they can be parsed and processed.
The competing desires are to model all language with the biggest tokens possible, and the fewest tokens possible. The lines aren't meaningless, text is split into the largest possible chunks using a set of the most common tokens.
Common words, like "the", "fast", "unity", "flying" are all tokens, but it's not because they're words, it's because they're common letter clusters, undistinguished from "fl", "ing", "un", "ple"
"gadflying" is tokenized into [g, ad, flying], even though it's only loosely semantically related to "flying", it's just the most efficient way to tokenize it.
1. Greatly reduces memory usage. Instead of memorizing every inflection of the word "walk", it memorizes the root (walk) and the modifiers (ing, ed, er, ...). These modifiers can be reused for other words.
2. Allows for word compositions that weren't in the training set. This is great for uncommon or new expressions like "googlification" or "unalive".
If you put:
test walk walker walking walked
into the tokenizer you will see the following tokens: [test][ walk][ walk][er][ walking][ walked]
Only walker is broken up into two different tokens.I added "test" to that because walk at the start doesn't include the leading space and [walk] and [ walk] are different tokens.
For even more fun, [walker] is a distinct token if it doesn't include the leading space.
test walker floorwalker foowalker
becomes: [test][ walk][er][ floor][walker][ fo][ow][alker]
How we think of words doesn't cleanly map to tokens.(Late edit)
walker floorwalker
becomes tokenized as: [walker][ floor][walker]
So in that case, they're the same token. It's curious how white space influences the word to token making.