Gigatoken: Fastest Tokenizer
twitter.com
twitter.com
I assume this tradeoff is purely for speed/compression. Or am I missing what's going on here?
There are many research papers on models using characters directly. One challenge is that effective context length is smaller.
Cool project nonetheless, I will go through the code later tomorrow