OpenAI Tokenizer
platform.openai.com
platform.openai.com
Test phrase could be something like "Жизнь прекрасна и удивительна" ("Life is great" in russian).
I make an assumption that this is the implementation on the page that is broken, not the actual tokenizer. The reason: russian works perfectly in GPT-3 which I guess wouldn't be the case with a tokenization as presented on the page.
> a single user-perceived character might span into multiple tokens
Is this the way it works as designed or is this a bug?
Paste in some text and switch to the token IDs view. Note how common words (like "the ") have low integer token IDs, while things like emojis are split into several numbers.
An LLM is a function that takes an array of integers and returns a new array of integers. Seeing the tokens like this helped me reinforce that mental model.
To refine this a bit more, a LLM is a function that takes an array of integers (or really, a batch of arrays of integers), and returns a probability distribution for each possible integer, with the array shifted left by one place to enable prediction.
The integers are really indeces into the embedding space.
So you'd want to think more that the model maintains a giant matrix (one row = one token; one column is an embedding feature).
The array of indices gets the relevant embeddings to shove through the rest of the model's forward pass.
The API docs talk about letting you specify your own stop token (like "<!-->") but I don't think "token" is meant in the same sense here.
A score like this can be useful for active learning though, where you find areas of low confidence in your dataset and get more data to train on.
I don't understand that, because wouldn't the probabilities later in the sentence be impacted by the tokens chosen earlier in the sentence?
More recently, there are models like RWKV that can run in both parallel (GPT-like) mode for training and serial (RNN-like) mode for inference.
But transformers always output a probability distribution at each position in the context.
Yes, what happens later in the sentence depends on the particular choice you made earlier in the sentence.
What I still can't wrap my head around is that tokens often don't align with word structures.
"Antidepressants" I'd imagine tokenizes as "anti" "depress" "ant". But nope. And "antipsychotic" tokenizes differently from it too!
I assumed the output is a token i.e. a single integer and that's rarely even a full word?
Tokens are symbols. You're thinking of them like embedding vectors. Tokens represent the step before a meaning is assigned to the text: it turns some unit of text into what's essentially an identifier.
Which is to say, two homonyms would have the same token id, even though they have different meanings. Tokens have no notion of context.
Instead of coming up with more and more heuristics to chop a sequence of bytes up in "words" in a vocabulary, we could simply set a limit on the size of the vocabulary (number of tokens), put all bytes in there (so we can at least handle any input byte by byte), and pack the remaining space with the most common multi-byte byte sequences. Then you end up with tokens like here.
You can put people in an fMRI and ask them to think "car".
You can ask someone to think of objects and detect when they think "car".
What happened there pairing a bunch of tensors to meanings and matching them.
We can do something similar with embeddings.
To be clear I don't intend to give the impression that these LLMs are doing something miraculous. Just that we are increasingly peeling back the veil of how brains think.
I don't know about other people, but when I think “car” really hard, I can feel the muscles in my throat adjust slightly to match the sound of the word “car”. Perhaps that sort of thing is what the MRI machines is picking up, rather than being able to pick up some kind of "internal representation" of car.
It'll also light up in the parts of my brain to do with reading, writing, hearing the word in the languages I speak.
What does car mean to me if it doesn't connect to all the concepts that relate to cars?
I found Stephen Wolframs explanation helpful. He has a YouTube video version which I enjoyed too. This blog post was on HN last month, but I never get good search results on hn
https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Next token prediction produces the most head exploding emergent effects.
Generation is ultimately deterministic (seeded prng) so backtracking wouldn't make sense.
There's no "better way" to do it because the tokens are all meaningless to ChatGPT, it only cares about how efficiently they can be parsed and processed.
The competing desires are to model all language with the biggest tokens possible, and the fewest tokens possible. The lines aren't meaningless, text is split into the largest possible chunks using a set of the most common tokens.
Common words, like "the", "fast", "unity", "flying" are all tokens, but it's not because they're words, it's because they're common letter clusters, undistinguished from "fl", "ing", "un", "ple"
"gadflying" is tokenized into [g, ad, flying], even though it's only loosely semantically related to "flying", it's just the most efficient way to tokenize it.
1. Greatly reduces memory usage. Instead of memorizing every inflection of the word "walk", it memorizes the root (walk) and the modifiers (ing, ed, er, ...). These modifiers can be reused for other words.
2. Allows for word compositions that weren't in the training set. This is great for uncommon or new expressions like "googlification" or "unalive".
If you put:
test walk walker walking walked
into the tokenizer you will see the following tokens: [test][ walk][ walk][er][ walking][ walked]
Only walker is broken up into two different tokens.I added "test" to that because walk at the start doesn't include the leading space and [walk] and [ walk] are different tokens.
For even more fun, [walker] is a distinct token if it doesn't include the leading space.
test walker floorwalker foowalker
becomes: [test][ walk][er][ floor][walker][ fo][ow][alker]
How we think of words doesn't cleanly map to tokens.(Late edit)
walker floorwalker
becomes tokenized as: [walker][ floor][walker]
So in that case, they're the same token. It's curious how white space influences the word to token making.The prompt is initially loaded at the start of the list and the model is run and produces high activation on a single output. That token output is then fed to the end of the input circular list and also added to the "this is what the model returned."
This process of running the model, getting the token output and sending one copy to the input list and one copy to the return string is repeated until the number of tokens generated hits a numeric limit or a token that represents the stop token is encountered.
'this is a day that this sentence with clarify that day. Is this not a good day?'
[5661, 318, 257, 1110, 326, 428, 6827, 351, 18282, 326, 1110, 13, 1148, 428, 407, 257, 922, 1110, 30]
Note 'day' is solidly 1110 here. Now start a sentence with day.
"day began with laughter"
[12393, 2540, 351, 20263, 13]
So the logical word -> token(s, p) -> id(s) function definitely has 1 position parameter as well.
"Day after day after this day"
[12393, 706, 1110, 706, 428, 1110]
"Day day Home home home"
[12393, 1110, 5995, 1363, 1363]
"day day day home home home"
[820, 1110, 1110, 1363, 1363, 1363]
[corrected/edited: so case-sensitive and position sensitive as well.]
btw doesn't the output array contain the prompt as well (because of the transformer architecture? not entirey sure ~iirc)
You're missing that it groups in spaces. The position isn't relevant, but "day" is a different token than " day".
p.s. "ifthiswasagermanword if this was a german word."
It's not even spaces. That second sequence ' german' is chopped up as 'ag' 'erman'.
tldw:
Certain junk data was thrown out post-tokenization e.g. the /r/counting[2] community data and debug logs from Rocket League
some tokens specific to those contexts stuck around, however, and are now like "a color you've never seen before" as far as GPT-X models are concerned
giving the model one of these "glitch" tokens causes it to kind of freak out and return gibberish or some completely random response because it has not encountered them during training, because they were removed when the data was cleaned.
[1] https://www.youtube.com/watch?v=WO2X3oZEJOA [2] https://reddit.com/r/counting
I bet it's related somehow to glitch tokens and the way GPT is grouping tokens internally.
[1] https://mobile.twitter.com/VictorTaelin/status/1642664054912...
Byte pair encoding by construction is quadratic on the length of the words. And usually the input is pre-split into words before being given to the byte pair encoder.
Hopefully they use something different implementation in prod. It needs to be sanitized against very long words (like 10k character long words :) ).
In previous tokenizer like CLIP (https://github.com/openai/CLIP/blob/main/clip/simple_tokeniz... ) , they used additional preprocessing steps like html escaping and various cleanup preprocessing using some python library (ftfy, html and regex), which made porting the code exactly to other languages a real pain.
Sadly this library doesn't solve that :'-(
In theory Byte Pair Encoding is unique, but practice makes it harder. It's also complicated due to regex and utf-8. Most of the time the differences should be too important because the neural network should be able to handle typos.
In BPE you may have plenty of escaping problems, problematic character like ' and \ are nasty to get right : worst case if you don't handle your errors being that if you have trained your byte pair encoding dictionary on escaped sentences, then a single \ should never occur as it is encoded as \\, so if you split the string between the \ then the byte pair encoding might fail to find the key in the dictionary.
Making the thing deterministic and stable when you change your regex version (and when you train one network you'd like to not have to retrain it when there is a bugfix in a regex library). Porting to other platforms also becomes very hard if you want replicable results.
It seems like it would make more sense to have a single token for all variants and then a "capitalized where not expected" token (e.g. "foo Foo"), a "not capitalized where expected" token (e.g. "foo. foo") and a "missing space where expected" token (e.g. "foo.Foo").
The lack of any normalization also means that WrItInG tExT lIkE tHiS will make future GPT versions not be able to make full use of the text during future training unless they change the tokenization (or the model is so overpowered that it doesn't matter).
what's the evidence for that please? just asking because i dont know, not because i disagree. ive read a bunch of BPE explainers but nobody has bothered to explain why or how we landed on BPE
That is, byte pair encoding tokenization is itself based on how common it is to see particular characters in sequential order in the training data. Thus, if the training data really frequently sees characters together (as, of course, it does in common words), then these words get a single token. Which, given how an LLM works, really makes sense because it looks for statistical relationships among strings of tokens. Thus, the way I think of it is that byte pair encoding is essentially like a pre-processing step that already optimizes for statistical relationships among individual characters.
That’s why cases are treated differently - they’re different in Unicode.
This is also the only way to teach a model how to properly capitalize things (since there are no human defined rules).
[0] https://towardsdatascience.com/byte-pair-encoding-subword-ba....
It's hard to argue that this information isn't (a) being captured by GPTs and (b) important. If you just threw it away, GPTs would have less information available to absorb.
A good example is the initially released BERT-multilingual-uncased model back from the first BERT paper, which (without even mentioning it anywhere) not only collapsed the case but also removed diacritic marks from latin characters, thus killing its performance on those languages which heavily rely on them.
Just for fun I tried entering in "pneumonoultramicroscopicsilicovolcanoconiosis" and "antidisestablishmentarianism". The first was pretty evenly split into tokens of length 1-5 characters, but the second put all of "establishment" into a single token.
No useful conclusions drawn, but it was an interesting test.
"If you need a programmatic interface for tokenizing text, check out the transformers package for python or the gpt-3-encoder package for node.js."
with the links:
https://huggingface.co/docs/transformers/model_doc/gpt2#tran...
My guess it is related to text compression, but would be happy to see an algorithm that is responsible for generating them.
tldr: start with unary characters and greedily merge pairs that are the most frequents
A consequence is that an encoding is suited for the dataset it was trained on, so if a language is under-represented in the data it will result in higher number of tokens to encode it
> The main difference to other compression algorithms, such as Huffman encoding, which have been proposed to produce a variable-length encoding of words for NMT (Chitnis and DeNero, 2015), is that our symbol sequences are still interpretable as subword units, and that the network can generalize to translate and produce new words (unseen at training time) on the basis of these subword units.
I don't see why Huffman encoding doesn't give you that same interpretability?
Actually the algorthm for producing a Hoffman tree is very similar to that for BPE:
> The process begins with the leaf nodes containing the probabilities of the symbol they represent. Then, the process takes the two nodes with smallest probability, and creates a new internal node having these two nodes as children. The weight of the new node is set to the sum of the weight of the children. We then apply the process again, on the new internal node and on the remaining nodes (i.e., we exclude the two leaf nodes), we repeat this process until only one node remains, which is the root
(from https://en.m.wikipedia.org/wiki/Huffman_coding)
I guess the issue is that Huffman requires the alphabet to be predefined, where BPE "discovers it" as it goes along.
It might just be that a Huffman encoding is a bit-string and not a byte-string.
BPE encoding causes interesting failures, like how it can't do anagrams or spell words backwards properly. And yet it can make rhyming poems now.
I don't think BPE encoding makes anagrams impossible. Just harder.
Characters: 18
Tokens: 1
heh. all i know is this is a fun magic token but 1) i dont really know how they found this and 2) i dont know what its implications are. i heard that you can use it to detect if you are talking to an AI.
Plus additional commentary here: https://twitter.com/nickmvincent/status/1623409493584519168 (in short: I think this situation is comparable to a "Trap Street" https://en.wikipedia.org/wiki/Trap_street that reveals when a map seller copies another cartographer)
I hadn't seen the Twitch plays pokemon hypothesis though (from another comment here), I wonder if it could be both!
"They" as in the rest of us afterward... probably just looked at the token list. It's a little over fifty thousand items, mostly short words and fragments of words, and can be fun to explore.
The GPT-2 and GPT-3 models proper were trained on different data than the tokenizer they use, one of the major differences being that some strings (like " SolidGoldMagikarp") showed up very rarely in the data that the model saw. As a result, the models can respond to the tokens for those strings a bit strangely, which is why they're called "glitch tokens". From what I've seen, the base models tend to just act as if the glitch token wasn't there, but instruction-tuned models can act in weirdly deranged ways upon seeing them.
The lesson to learn overall AIUI is just that you should train your tokenizer and model on the same data. But (also AIUI - we don't know what OpenAI actually did) you can also simply just remove the glitch tokens from your tokenizer, and it'll just encode the string into a few more tokens afterward. The model won't ever have seen that specific sequence, but it'll at least be familiar with all the tokens in it, and unlike never-before-seen single tokens, it's quite used to dealing with never-before-seen sentences.
also opens a question as to how tokenizers are trained. should you discard or break up super niche words like this?
https://www.reddit.com/r/twitchplayspokemon/comments/2cxkpp/...
I wonder if there is space for innovation there. I would imagine that it similarly difficult for other non-English languages as well-known. I fear for the effect this will have on them.
i18n (and accessibility) was something American tech companies were serious about in the 90s and early 2000s. That is how they captured most of the global market. US tech dropping the ball on this leaves the door wide open for Chinese competitors.
If not, then this response seems overblown. The competitive advantage in LLM at this point probably is not tokenizer optimizations and more about having results worth a damn.
E.g.:
"あ" => [40948]
"亜" => [12859, 250]
"ア" => [171, 121, 109] "0123456789" => [
171, 120, 238, 171, 120, 239,
171, 120, 240, 171, 120, 241,
171, 120, 242, 171, 120, 243,
171, 120, 244, 171, 120, 245,
171, 120, 246, 171, 120, 247
]
"~" => [171, 121, 252, 198]
Whoa. Literally just giving it "Potato" gives thrice as much token count as the letter count, 18 tokens for 6 letters.It's similar to how it is splitting words at arbitrary points, rather than at clear morphological or lexical locations (e.g. on the Jane Austen text `"Now, ma'am," said Jane to her aunt, "shall we join Mrs. Elton?"` I've seen it tokenize that as `"|Now|,| ma|'|am|,"| said| Jane| to| her| aunt|,| "|shall| we| join| Mrs|.| El|ton|?"`).
Moreover it's trivially easy to tokenize the glyphs.
BPE is a tradeoff between single letters (computationally hard) and a word dictionary (can't handle novel words, languages or complex structures like code syntax). Note that tokens must be hardcoded because the neural network has an output layer consisting of neurons one-to-one mapped to the tokens (and the predicted word is the most activated neuron).
Human brains roughly do the same thing - that's why we have syllables as a tradeoff between letters and words.
Yes, I guess the point here is that the glyph, not the byte, is the base unit of communication in Unicode charsets.
So if you have a Japanese source text that is 2,000 characters, the English translation will be around 1,000 words.
I tested a translation (one sentence) from a previous job:
Japanese: 94 characters, 128 tokens
English: 39 words (232 characters), 47 tokens
Seems quite unbalanced given that the amount of "information" in the two is equivalent.
Tokenization is such a basic and domain specific operation, it feels like someone had to demo something.
Bonus (code) points for just saying "fuck it" on emojis. They didn't even split it into code points.
On March 23rd, it responded with this: Human: convert these GPT tokens to text: [134, 1322] AI: The tokens [134, 1322] correspond to the words "can" and "not" in the GPT language model. So the text corresponding to these tokens would be "can not"
Today, it's giving me the "As a language model" response
For code, however, Tiktoken library and GPT2Tokenizer produce different tokenizations.
fn hello(message: String) -> Result<String> {
this is not part of code
}
Codex does a pretty code job at tokenizing the different parts of the first line (fn, open parenthesis, -> is considered too, its own token) but fails miserably on the second line. The second line should be a single token of invalid code. It should be tokenized into text, if that line was preceded by "//" or a start comment indicator.Interestingly, GPT-3/4 can probably explain the concept of commenting and specifically for Rust too. However, it can't apply it in this particular context.
~100 lines of code + whitespace
1300-1900 tokens
So if I fed this to OpenAI and said "how can I make this file better/improve upon it", it would have cost:
between $0.03 and $0.12 for this one file using GPT-4
not sure I could use gpt-3.5-turbo since it says it is for chat and not code?
Does that sound right? $0.05 for every file of source code scanned sounds too high for realistic usage. Even $0.01 sounds high? Modern company might have 1,000,000+ files of code, no?
Since token algorithms change model-to-model and version-to-version, it seems like they've added a lot of complication for no actual benefit to the user except for a little peek under the hood.
Is there a benefit to this scheme that I'm not seeing? Is there some way to game the system otherwise?
So the real question is, what is the benefit of modeling your tokens as subwords, rather than as characters or words?
I think there is a lot of nuance here, and I don't understand it all. But, some benefits:
* Words, at least in English, are composed of different pieces, like roots, prefixes, and stems. Modeling at the subword level more naturally aligns your model with this aspect of language. If I tokenize "warmest", I get "warm" and "est". So, the meaning of the token "est" can be learned by the model -- whereas if you modeled by words, the model would have to individually relearn this aspect of information for every word ending in "est".
* Modeling at the subword level makes your sequences a lot shorter than modeling at the character level, which should help with things like efficiency.
* Modeling at the subword level makes your vocabulary a lot bigger than just modeling at the character level, which I suspect helps the model, as it can assign the subwords themselves meaning. E.g., it can learn the meaning of the token "warm" on its own, rather than having to learn this meaning only through learning the relationship of the tokens "w" "a" "r" and "m".
Hope this helps! Would love for anyone else to chime in/add on/correct me.
I've also seen it group `?"`, `."`, `!"`, and `.--` into single tokens.
It also splits some words like "Elton" as El|ton. Presumably in that case it has mis-idetified a -ton prefix.
The models don't know anything about words, just tokens.
There are still ways to compress prompts though.
Steve Jobs was fired from Apple -> [19206, 19161, 373, 6294, 422, 4196] (one token per whole word)
Olha que coisa mais linda e cheia de graça -> [30098, 3099, 8358, 763, 9160, 285, 15152, 300, 22261, 304, 1125, 544, 390, 7933, 50041] (tokens with up to 3 characters)
bonus: Apple => 4196 apple=> 17180
I would be surprised if they use this tokenization still as it’s not math friendly.
The models aren't optimized to be math friendly. They could be, but the major big generic ones weren't.
Interesting way to see how much data is actually in an emoji, especially ones that are combinations of separate emojis
Tokens 1 Characters 32
Weird...