Aren’t there only 2^16 tokens? Seems easy to test for all of them, but I might just not understand the tokenizer.
From what I've found through Google (with no real understanding of llm) 2^16 is the max tokens per minute for fine tuning OpenAI's models via their platform. I don't believe this is the same as the training token count.
Then there's the context token limit, which is 16k for 3.5 turbo, but I don't think that's relevant here.
Though somebody please tell me why I'm wrong, I'm still trying to wrap my head around the training side.
If you have any other resources (for anything AI related) please share!
Data files that contain vocabulary are listed here: https://github.com/openai/tiktoken/blob/9e79899bc248d5313c7d...
[1] https://github.com/openai/tiktoken/blob/main/tiktoken/model....