LongRoPE: Extending LLM Context Window Beyond 2M Tokens
arxiv.org
arxiv.org
I know people complain about hardware and compute resources. But, this is like complaining about Python resource usage in the early 90's. The development complexity & resources is far more expensive than chips on the long run. I am personally re-organizing my AI organization to move away from complex RAG setups and get comfortable with long-context workflows.
Just to be clear - I also think that inference-optimized chips are the next frontier - Nvidia GPUs were designed & built in a different age than what’s going on now.
It's really not even true for your python example.
E.g. can produce accurate scene-by-scene descriptions of hour long movies, can answer questions about very large codebases, etc.
I can see Q&A on a movie being useful but can't think of descriptions of the overall movie itself lacking... I'd definitely be worried about spoilers.
The current landscape is moving so fast that we don't stop and work on achieving the best version of a specific approach. One can argue that both RAG and having 1M+ context length is a symptom of a disconnect between what language models are capable of and what people think they can achieve with them.
Within the next 6 months there might be yet another approach to provide LLMs custom memory, and we will be having the same discussion the way things are handled currently.
Because paying like $100 for a 2M query wouldn't be cool for most of us, so the applications would be very limited.
I'm not as familiar with CPUs as I am with mathematical concepts. I don't know what the name for the processor bit hacking tricks is called. But that's maybe the general idea for data compression for LLMs/transformer models on CPUs, I think.
After all, notice how data compression improvements are only multiples of two. 128k tokens and 2048k tokens. There's an implementation dependent CPU optimization hack going on in there somewhere.
I'm a bit out of my depth, but I think ultra-long exact-attention work like this also probably has to answer some questions about where to put the KV-cache before it can be used in practice?
But since this work depends on strong pre-trained model to extend from, I think it's open whether a training a byte-level model from scratch with similar tricks would result in the same performance (and whether any organization in the world has the GPUs and chutzpah to do pre-training at the long context lengths...)
[0]: https://arxiv.org/abs/2305.07185 [1]: https://arxiv.org/abs/2401.13660
- make vocab ~40k tokens + 256 tokens (for each possible byte),
- start training with 100% tokens,
- then after some portion of train budget has elapsed, randomly replace some tokens with corresponding byte-tokens,
- then ramp the fraction of replacements up so you're training on 100% byte tokens for the last xx% of training, without ever exceeding an 8k (or whatever) sequence length
- then apply the trick from TFA to get xk -> 8xk/64xk bytes of context?
but I'd guess the interesting part of a byte-transformer is multimodality, and we'd need more than a few tricks to get from ^^ to there.
For what it's worth, we actually were working with LSTMs with nearly a billion params back in 2016-2017 area. Transformers made it far more effective to train and execute, but ultimately LSTMs are able to achieve similar results, though slow & require more training data.
Why? I think this is because books3 contains, you know, books - including some really long books, and proof-pile contains math papers and math stuff, which isn't as long.
So overall I think what you're seeing is a general trend of increasing perplexity on windows above 256k, between 256k-2048k, which is probably not so surprising - or at least, not so surprising when you consider the context of the paper, which is taking a model pre-trained with a much shorter context window and extending the context window using a novel technique. It's hard to adapt a model trained to do one thing into doing another thing, and that's what they're doing, so in that context, it tracks that the longer the context window, the worse the performance.
Clearly hardware vendors love LLMs. But it's just a highly inefficient approach for a lot of problems.
As someone who's run into problems related to short context windows quite often, can you explain what "worth it" means? Also, when you say "not cheap", what alternative do you have in mind?
Of course it is true for real long context but it's not clear if its going to work with the sparse context hacks intended to keep memory size low.
Having the ability to keep the same coding session forever without having the LLM forgetting all the time and making the same mistakes over and over would be a game changer.
People like RAG-based solutions because you can include references in your answer very easily (eg, Perplexity, or see the DAnswer "internal search" product launched today on HN). That is extremely hard to make work reliably from a fine-tuned model.
They often want to see the reference - for example the LLM is constructing an answer about some company policy question it is good to include references to the actual company policy documents by URL. This is easy using RAG, but very hard to do reliably just using a fine tuned LLM.
See Perplexity for a great example of this at web scale.
But nonetheless, sending 2M tokens to a LLM isn't an efficient way to solve most problems.
I think there is sufficient evidence to think it works. A bunch of people who do have access have put it through a bunch of real-world exercises to test this, and while GPT4 128K and Claude 200K both fail Gemini is passing.
I agree sending 2M tokens to a LLM is costly at the moment and that RAG has a bunch of advantages in some circumstances.
Figuring out new algorithms is monumentally more difficult than more hardware, and even then when more efficient algorithms are found, we throw more hardware at it and get a million times more done.
However, caching might be a sweet spot for these multi-modal and large context LLMs. Take a bunch of documents and perform reasoning tasks to distill the knowledge down into something like a knowledge graph, to be used in RAG.
They would still be valuable for prototyping, where fast iteration makes it possible to learn more about the problem you are solving and whether it is even worth solving. They are also valuable for iteration-and-distillation approaches, where you can use the data generated from an expensive model to train a cheaper model.
So it looks like a VERY good idea. Who gets the gist of what I'm saying?
You'd get a repeat of Microsoft Tay, the short-lived 2016 chatbot that had to be taken down after not even a full day because 4chan managed to turn it into a full-blown hate spreader in hours [1].
Before something like this can be reasonably released to the Internet at large, we need to find out how to teach the "ingestion" part how to judge the input that it's being presented with... basically, AI school. The same way we teach our children that it's not OK to steal other people's stuff, to not be the one first throwing punches or to not discriminate against other people, or that there are reasonably trustworthy media and absolutely untrustworthy, we need to teach AIs. And we'd also need to figure out ways to teach an AI basic, hard truths: the Holocaust happened, the Earth is a globe not a pizza, the Earth is not hollow, and the moon landings were real.
At the moment, we're half-ass attempting that by annotated training data (and this is the true moat of OpenAI, not the weights or prompts!), but even as we invest literally millions of hours of training compute time, an average high schooler that has been trained for 18 years can reasonably pass the above-mentioned criteria to be a productive member of society.
This is precisely why programs that are nourished with innovative models, backed by substantial computational power, are capable of developing reasoning akin to Q*. Once these models start to independently foster these advancements and self-improve, we'll witness an unprecedented surge in AI development, surpassing our current capabilities.