You don't have to use tiktoken if you aren't actually tokenizing things. The token lists are just text files that consist of the characters base64 encoded followed by the numeric ID. If you want to explore the list you can just download them and decode them yourself.
I find that sorting tokens by length makes it a bit easier to get a feel for what's in there.
GPT-4 has a token vocabulary about twice the size of GPT-3.5.
The most interesting thing to me about the GPT-4 token list is how dominated it is by non-natural languages. It's not as simple as English tokenizing more efficiently than Spanish because of frequency. The most common language after English is code. A huge number of tokens are allocated to even not very common things found in code, like "ValidateAntiForgeryToken" or "_InternalArray". From eyeballing the list I'd guess about half the tokens seem to be from source code.
My guess is that it's not a coincidence that GPT-4 both trained on a lot of code and is also the leading model. I suspect we're going to discover at some point, or maybe OpenAI already did, that training on code isn't just a neat trick to get an LLM that can knock out scripts. Maybe it's fundamentally useful to train the model to reason logically and think clearly. The highly structured and unambiguous yet also complex thought that code represents is probably a great way for the model to really level up its thought processes. Ilya Sutskever mentioned in an interview that one of the bottlenecks they face on training something smarter than GPT-4 is getting access to "more complex thought". If this is true then it's possible the Microsoft collaboration will prove an enduring competitive advantage for OpenAI, as it gives them access to the bulk GitHub corpus which is probably quite hard to scrape otherwise.