You can use their online tool to see how it tokenizes words: https://platform.openai.com/tokenizer
You can test it offline using tiktoken: https://github.com/openai/tiktoken
The tokenization speedups in that repo are very impressive. It was the most annoying part about processing 190,000 books. I think it took a few days on a server with 96 cores.
Surprisingly hard to figure out the vocab size from that repo.
The vocab size itself is doubled. (~50k for GPT-2/3, ~100k for ChatGPT)
It certainly makes training more expensive. One clever trick to get some memory savings is to freeze the vocab embedding layer when fine tuning. It makes a noticeable improvement, both in speed and in mem required.
Surprised they went the larger vocab route. LLaMA is only 30k. I wonder what the reason is...
Thanks!