8192 words is getting into the range of short stories or a masters thesis, which opens the door to some interesting applications.
8192 words is getting into the range of short stories or a masters thesis, which opens the door to some interesting applications.
That said, your point stands. Most short stories are low-to-mid four-digit words, and a jump from 2048 tokens to 8192 squarely fits in that window.
As someone who's been working on multi-layered approaches to using GPT-like models for long text generation (e.g. synopsis -> outline -> paragraph expansions) to get around the limited context window, it'll be interesting to see if people will keep working towards that end or if it'll all become a moot point as the effective context window continues to scale up.
Never mind, it looks like they have a tokenizer tool online and every bigram I’ve given it BPEs to multiple tokens:
It's unclear what tokenizer they are using and the documentation is being coy about it. It could be a more efficient or a less efficient tokenizer.
The code there implies cl100k_base has a vocab size of 100k (I guess it's in the name lol) which means it is more comprehensive than GPT-2's 50k, so fewer tokens will be necessary.