"Studying the limits of Gemini 1.5 Pro's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens"
https://storage.googleapis.com/deepmind-media/gemini/gemini_...
"Studying the limits of Gemini 1.5 Pro's long-context ability, we find continued improvement in next-token prediction and near-perfect retrieval (>99%) up to at least 10M tokens"
https://storage.googleapis.com/deepmind-media/gemini/gemini_...
Having 99% retrieval is nuts too. Models tend to unwind pretty badly as the context (tokens) grows.
Put these together and you are getting into the territory of dumping all your company documents, or all your departments documents into a single GPT (or whatever google will call it) and everyone working with that. Wild.
Input is parsed one token at a time right? Can you cache the state after the initial prompt has been provided?
>Finally, we highlight surprising new capabilities of large language models at the frontier; when given a grammar manual for Kalamang, a language with fewer than 200 speakers worldwide, the model learns to translate English to Kalamang at a similar level to a person learning from the same content.
Results - https://imgur.com/a/qXcVNOM
But there is also the downside of "tuning" the RAG to return less tokens you will miss extra context that could be useful to the model.
One point I'm unclear on is how these huge context sizes are implemented by the various models. Are any of them the actual raw "width of the model" that is propagated through it, or are these all hierarchical summarization and chunk embedding index lookup type tricks?
The Encyclopedia Britannica is ~44M words.
As the other comment mentions, you can paste the content of entire books or documents and ask very pointed question about it. Last year, Anthropic was showing off their 100K context window, and that's exactly what they did, they gave it the content of The Great Gatsby and asked it questions about specific lines of the book.
Similarly, imagine giving it hundreds of documents and asking it to spot some specific detail in there.
It's also not obvious how these huge models will fare against increasingly capable open source ones like Mixtral, perhaps especially since Google are confirming here that MoE is the path forward, which perhaps helps limit how big these models need to be.