Tiktoken: OpenAI’s Tokenizer
github.com
github.com
* the cl100k_base tokenizer has ~100k tokens -- previous tokenizers had ~50k. (enc.n_vocab gives 100277 but some numbers in that range don't work, starting at 100256)
* it has exactly 1110 tokens which are just digits. 10 1 digit tokens, 100 2 digit tokens and 1000 3 digit tokens! (none have preceding spaces). this is a huge improvement from GPT2's tokenizer, which was a huge mess.
* there are <|fim_prefix|>, <|fim_middle|>, and <|fim_suffix|> tokens (see Efficient Training of Language Models to Fill in the Middle)
The biggest news to me is the improved handling of numbers. This could explain some improved performance on arithmetic. One disappointment is that it tokenizes from the front, e.g. "1000000" -> 100|000|0. This is one of those "so close!" moments -- I would work for free to fix this.
In that context "open sourcing many useful projects" seems the bare minimum it should be doing, because that was the promise still enshrined in it's name - this is not a bog-standard commercial organisation where that would not be expected.
To be clear OpenAI does deserve massive kudos for it's achievements but openness is not among them.
I hope pypi libraries can provide complete standalone offline versions instead of requests+urllib3+some_object_storage shenanigans.
If these blobs are too large to host it on pypi, maybe give us an alternative way to download it altogether so we can deploy the full lib to a server without network access?
What's the point of this MIT license, then: https://github.com/openai/tiktoken/blob/main/LICENSE
....going closed source and monetizing.
1. Name
2. OpenAI is releasing useful stuff
3. Rust in AI!
- javascript: https://www.npmjs.com/package/gpt-3-encoder
- c# https://github.com/dluc/openai-tools
- java: https://www.reddit.com/r/MachineLearning/comments/upej7e/p_j...
- php: https://github.com/CodeRevolutionPlugins/GPT-3-Encoder-PHP
"gpt2": gpt2,
"r50k_base": r50k_base,
"p50k_base": p50k_base,
"cl100k_base": cl100k_base,Or maybe they do a quick pass on websites and filter out websites with bad words quickly, just leaving a small fraction for the model to train on. Parsing to get the standardized tokens quickly would make that easier.
sample = "ฟองมันฟันหนู, ฟันหนูฟองมัน, ฝนทองฟองมัน"
len(sample), len(enc.encode(sample))
This returns `39, 40` so it's just encoding one character at a time. It's probably like this for almost all non-English text.Tokens: [split, an, input, into, tokens, .]
Tokens: [umb, rella] ?
This particular tokenizer is very interesting given that it tries to be best of both worlds (word-level tokenizer and character-level).
(Could we say it is kind of a garden-path word?)
facebook-ed: Includes the whole name, but nothing else. Obviously bad and not OK.
tik-token: Includes the whole name. Also includes a generic common noun that relates to the product at hand, and is completely fine to include in the name. Appart from that "tik" can not be said to "imply a connection to tiktok". This is fine.
More serious "qiktoken" would at least hint at the use for the package.
For regex specifically, the usual "regex" crate is fantastic and provides very good performance guarantees (no exponential behaviour, ever!), unlike a lot of regex libraries. However, it also comes with the tradeoff of not supporting arbitrary lookahead/lookbehind or many regex extensions, so people needing those features might need a different crate.
There's a nice discussion on another thread: https://news.ycombinator.com/item?id=33510976
In Javascript this led to, well, npm. Because there's no standard library, and everyone ended up reinventing everything.
It's easy to find drawbacks of either extreme. I find Rust's approach of favoring competing, versioned libraries over standard library inclusion generally works well. Many of the libraries managed to significantly improve performance and ergonomics because they can have space to improve their API, and explore which behaviors and guarantees are useful. The biggest downside is that it's more difficult to protect against supply chain attacks.
Supply chains are real and being scared to design good API that will survive test of time sounds like poor excuse
At worst you could steal API design from some mature and battle tested base class libs like .net
There is nothing stopping us from having a small standard library and a set of official - but optional - crates for common things like regex. This is the best of both worlds in my view.
- You type something in expecting autocomplete, its not there. "Huh, thats weird"
- Look it up online. Sure enough, its not in the standard library. OK, what should I use then?
- Do research to figure out what 3rd party crate to use
- Tend to abstract/isolate those bits in case you need to rip it out for a different implementation later
To me the right approach is to publish an "official" meta package that depends on a few core packages that work well together.
Sure it might be overkill or "underkill ?) for some ?
Surely the output of said tokenizer you're going to stick through a machine learning model that is orders of magnitude slower?
So it really doesn't matter if the tokenizer takes microseconds or milliseconds, when the main model takes seconds for the same input/output.
And I believe the ratio is that big - you just use a lot of compute for training, and that's why it's so expensive.