From deep to long learning?
hazyresearch.stanford.edu
hazyresearch.stanford.edu
There are other interesting ongoing efforts to increase sequence length. One that has worked for me is this dynamic routing algorithm, related to self-attention, that can handle sequences with 1M+ tokens in a single GPU: https://github.com/glassroom/heinsen_routing . Right now, you can take 1,000 sequences of hidden states computed by a pretrained transformer, each sequence with, say, 1024 tokens, concatenate them into a single ultra-long sequence with 1,024,000 hidden states, slap 1,024,000 position encodings on top, and feed the whole thing to that routing algorithm to predict the next token (or whatever other training objective you want to optimize for). It works. Search the README for "Very Long Sequences".
If anyone here has other suggestions for working with long sequences (hundreds of thousands to millions of tokens), I'd love to learn about them.
Attention isn't really O(n2) anymore.
https://news.ycombinator.com/item?id=35375616
So far, every method I've come across that works well with long sequences does it by explicitly modeling (approximating) all possible pairwise relationships between tokens. In the case of attention, that means modeling n×n, or n², pairwise relationships. In the case of that dynamic routing method, it means modeling n×m pairwise relationships, where n and m are the number of tokens in the input and output sequences. I suspect that the key reason why that routing method works for long sentences is because m can be much smaller than n: Routing 1M tokens to 100 tokens requires ~10,000x less compute than applying attention over 1M tokens for each of 1M tokens.
If the work described by the OP succeeds in bringing computational cost down from O(n²) to O(n log n) while still explicitly modeling (approximating) all n² pairwise relationships, that would be a huge win.
- "Hyena Hierarchy: Towards Larger Convolutional Language Models" https://arxiv.org/pdf/2302.10866.pdf
- "Resurrecting Recurrent Neural Networks for Long Sequences" https://arxiv.org/pdf/2303.06349.pdf
It's unclear (at least to me) if any of these scale as well as transformers at this point.
I don't think Hyena can be run in RNN mode, though. But it _might_ scale better than the LRU model (second paper).
May I ask what you are using it for?
I've found it to be a useful tool. It's become my go-to solution whenever I need to combine data, hidden states, or model outputs of different sizes/dimensions/modalities into something that works on the first try. It can "glue models together into systems" without much hassle. I mean, you still have to tinker and tweak to get things to work well, but no more than usual.
Most recently, I've been using it to build models with access to long-term memory, more or less like this: Concatenate a large number of sequences of previously computed transformer hidden states. Apply a couple of routing layers to shrink them into a single short sequence of fixed size. Attach the shrunk sequence as a form of long-term memory to the current sequence of final transformer hidden states. Slap a couple of routing layers on top to predict things of interest. Finally, train only the routing layers (leave all the transformer stuff frozen). You can of course save the long sequences of previous transformer hidden states in a key-value store and fetch them as needed :-)
If you're looking to make context longer, a more practical approach is to route multiple past sequences of frozen pretrained transformer hidden states to a much shorter sequence, say, of length C, concatenate that with the current final sequence of hidden states, say, of length N, apply masked autoregressive attention on the C+N hidden states and then a feedforward layer to predict all next tokens in parallel, and last, drop the first C predictions, i.e., use only the trailing N predicted next tokens for training and inference.
I think we're going to see a lot more in the wavelet/convolution/fft space when thinking about how to increase context length.
I think there's also a lot of room for innovation in the positional encoding and how it's represented in transformer models, it seems like people have been trying lots of things and going with what works, but most of it is like: "look, a new orthonormal basis!".
Hyena sort of seems like the first step in moving to positional embeddings (or joint positional/attentional embeddings).
Very cool work.
The main issue I've seen with other wannabe-sub-quadratic-replacements for self-attention is that they all rely on some kind of low-rank/sparse approximation that in practice renders LLMs incapable of modeling enough pairwise relationships between tokens to achieve state-of-the-art performance.
I'm curious to see if this kind of approach solves the issue.
That said, my sense is there's a significant difference between (a) learning to approximate the matrix of n×n interactions with more run-of-the-mill methods, such as some kind of linear decomposition with learned fixed coefficients, versus (b) learning to approximate the matrix of n×n interactions with dynamic methods that "find the most suitable decomposition for each sample on-the-fly," which is what these class of models appears to be doing.
Apologies if all this sounds very hand-wavy; it's the best I can do at the moment.
Off topic, but I am curious what hardware Apple will release in the future for more direct AI support. Their Core ML libraries working with Apple Silicon have been very effective so far. The next step would likely be a built in foundation LLM model, extending what they have supported with BERT models, etc.
I'm thinking deeper. It wouldn't surprise me if self-attention itself becomes a primitive building block of future co-processors, e.g., with instructions and memory layouts engineered to make ultra-low-precision self-attention as compute- and memory-efficient as possible. I'm expecting LLMs with hundreds of billions and eventually trillions of parameters will be able to run locally on my laptop and mobile phone, in the not-too-distant future.[a]
[a] If this sounds far-fetched, consider that you can already run LLMs with tens of billions of parameters on mobile phones: https://justine.lol/mmap/
Perhaps. There's been a lot of focus on training-compute optimal models in the industry. Rightfully so, as proofs of concept. That's what led to this perceived parameter count race in published models.
But remember the other side of the scaling laws. For inference, which is what we want to do on our phones, it's better to be inference-compute optimal. That means smaller models trained for longer.
As far as we know today there are no limits of the scaling laws. A 1B parameter model _can_ beat a 1T parameter model, if trained for long enough. Of course it's exponential, so you'd have to pour incalculable training resources into such an extreme example. But I find these extreme examples elucidating.
My pet theory these days is that we'll discover some way of "simulating" multiple parameters from one stored parameter. We know that training-compute optimal models are extremely over-parameterized. So it isn't the raw capacity of the model that's important. It seems like during training the degrees of freedom is what allows larger models to be more sample efficient. If we can find a cheap way of having one parameter simulate multiple degrees of freedom, it will likely give us the ability to gain the advantages of larger models during training, without the inference costs later.
I don't disagree that we're likely to see more and more parameter capacity from our devices. I'm just pointing out that the parameter count race is a bit of an illusion. OpenAI discovered the scaling laws and needed a proof of concept. If they could show AI reaching X threshold first, they could capture the market. The fastest way to do that is to be training-compute optimal. So they had to scale to 175B parameters or more. Now that it's proven, and that there's a market for such an AI, their and other's focus can be on inference-optimal models which are smaller but just as smart.
Good point. That could very well be what they're thinking about, in addition to potential improvements in training data and RLHF methods.
Also, I agree it would be great if anyone figures out how to do something akin to "making a smaller model act as if it were gigantic during training" OR "pruning a gigantic model's 'dead paths' as it learns during training," to get the benefits of scale in training without its costs at inference.
Although unfortunately I don't really understand any of it. But I strongly suspect that things need to be better factored if we want good interpretability and efficiency. Not that I think it's easy to do that and still get performance and generality.
But I do see a Neural Engine 2.0 in the future that will better handle these things in the more near term future.
Regardless, for most programming tasks, I doubt my equivalent human context length is any better than 32k tokens.
ChatGPT has already been hooked up to Wolfram Alpha.[1]
For hooking up other things, see HuggingGPT and TaskMatrix.AI.[2][3]
[1] - https://writings.stephenwolfram.com/2023/03/chatgpt-gets-its...
The fact that this is such a multi-dimensional problem is frequently overlooked in the debate about AI/AGI etc, it may not matter all that much if the non AGI AI is already superhuman on enough dimensions other than the ones that the 'but it isn't AGI' crowd cling to. The consequences are what matters, not the fine print or the implementation details, and those consequences are directly tied to the number of dimensions along which a computer can beat humanity.
To give an example: if a chess program was 3000 Elo before but so slow that it would lose under competition rules then humans still dominated chess. Likewise if it would be only 2500 Elo but fast enough, it would still lose from the best humans. But for a large fraction of society it would have already moved out into 'superhuman' territory. And a couple of technological leaps of progress later and we're all looking at that AI as if it has moved into superhuman regions.
This sort of thing will happen on many fronts, and all of those fronts are moving, if enough of them go past the threshold then whether it is AGI or not is irrelevant and for every person that threshold is at different points. Maybe a computer will be able to calculate faster and better than you can, maybe it will be able to translate text faster and better than you, maybe it will be able to organize information faster and better than you. At which point we cross the line into saying that it can think faster and better than you is hard, but we can see that we are getting close to that line without even knowing exactly where that line is.
What is considered human is malleable. It is conceivable that humans will be enhanced in various biological and non-biological ways to a point that they can once again compete with computers.
"Chess" is a human created game after all.
But moving goalposts is an integral part of all professional sports. The 3 point line in basketball and engine size and aspiration in motorsports are obvious examples. I don't see how adjusting chess rules to dis-favor AI competitors is any different.
How are you working out that it has a 3000 Elo if it's not winning games?
Stockfish has an Elo rating over 3500 in spite of no human being even close to that.
When you play a turn-based game with a bot, that means you don’t need to worry about its reaction time. It’s paused most of the time, waiting on you. A sorcerer’s apprentice scenario isn’t going to happen when you’re single-stepping.
Moving to routine use of bots that run continuously with fast reaction times will be much more dangerous.
I don't think this was ever true. Chess programs appeared EXTREMELY early in, and everyone recognised that it was a matter of time until hardware was quick enough to evaluate so many positions per second that grandmasters could be defeated by sheer calculation.
Humans have multiple layers of memory and can recall things and concepts from years in the past - akin to millions of tokens' worth of recall in an LLM. Yes, that memory is extremely lossy, but it's there.
Can someone explain to me why a transformer or RNN is better at this than a simple linear layer with an equivalent number of parameters? A linear layer can receive the context [1, 2, 3, 4, 5, 6, 7, 8], properly one-hot encoded / embedded, etc, and predict the next sequence. Can a linear layer do just as well as a transformer? This setup allows linear layers to predict sequences with an arbitrary context size, so why so much hype about transformers and RNNs and other sequence focused architectures?
Perhaps the difference is that given the same number of parameters, the transformer uses those parameters to perform easy computations whereas the linear layer just does one gigantic matrix multiplication which isn't very efficient?
I understand each context input is embedded with its position, but I suppose the transformer can learn to ignore the position and just look at the context as an unordered set?
https://arxiv.org/abs/2205.13504
from the abstract:
"Recently, there has been a surge of Transformer-based solutions for the long-term time series forecasting (LTSF) task. Despite the growing performance over the past few years, we question the validity of this line of research in this work. Specifically, Transformers is arguably the most successful solution to extract the semantic correlations among the elements in a long sequence. However, in time series modeling, we are to extract the temporal relations in an ordered set of continuous points. While employing positional encoding and using tokens to embed sub-series in Transformers facilitate preserving some ordering information, the nature of the permutation-invariant self-attention mechanism inevitably results in temporal information loss..."
Similar to what Numenta HTM networks do, but scalable and performant for real use cases.
BTW, perhaps human-like conscience emerge as a "self-attention-like" mechanism between context and learning. Just saying.
What I think (and this is just me talking out of my ass) will be required is some form of associative long-term memory. Basically, give the model a way to store some embeddings in some form of memory, and then retrieve them based on context: so it doesn't matter if you encountered that item 2 tokens ago, or 2B.
At least this is what my current intuition tells me.
Have there been any state-space models adapted for arbitrary text generation?
Language models like ChatGPT are trained to predict new words based on the previous ones and are excellent for generation, a harder task than translation or classification. I'm doubtful about the adaptability of text models that deal with fixed-sized input/outputs and don't have an architecture that is as natural for generating indefinitely long sequences.
As a pleb who doesn't even own a data center, I've been hoping that a superior machine learning architecture will be discovered that doesn't scale well. We would be fortunate if our personal computers end up being half as good as Microsoft's or Amazon's best models; fortunate if the best architecture gains little from an additional 10,000 GPUs. This would help spread the benefits of AI evenly among anyone with a phone or computer -- a utopia compared to the other possibility, that everyone can learn how to build AI, but only those with a few hundred million to throw at a data center can actually control the means of production -- err, I mean, the means of intelligence.
Philosophically, this wouldn't be unlike people. Humans are still the greatest intelligence we're aware of, and humans don't scale. I'm hoping computer intelligence ends up not scaling well either.
The (depthwise) convolutional realization is extremely efficient for training, and the RNN is extremely efficient for inference. The scaling in both of these cases is much better than attention layers - as they discuss in the article.
I think that they are following the easy path, much like the Giga Hertz race in CPU, and will hit a wall. Maybe the wall will be so far away that it will give us an AGI but maybe it will give us superhuman machines only in well defined contexts. We'll have to squeeze our instructions in a prompt too small for some tasks and get a bot behaving like the main character of the Memento movie (he remembered only the last few minutes and very old memories.)
Or in other words, it was always a new form of search.
Wait, it’s all just search?
[astronaut with gun]
Always has been
The model architectures are different, and in the very latest paper they scale these not-transformer models to sequence length of 64k, where the paper you linked only considers up to 8k
I disagreed with him, and this article is evidence that is in favor of my point. If research like this continues to move forward LLMs will improve at a rapid rate.
Different threads attract different groups of people with different areas of expertise. So I will sort of reiterate this topic here as I'm interested. What are most people's thoughts on this "local maximum" thing have we actually hit a dead end? Especially given the proliferation of effort towards producing research like the one showed in the topic here.
The researchers are also, excited.
We’re especially motivated by applications that could benefit from longer-sequence models – high-resolution imaging, new modalities of data, language models that can read entire books. Imagine giving a language model an entire book and having it summarize the plot, or conditioning a code generation model on all the code you’ve ever written. The possibilities are wild – and we’re excited.
As with most technology there isn't necessarily always a constant influx of inflection points and paradigm shifts. Improvement will likely creep up on us incrementally. Suddenly one day its clearly more intelligent then a human and we can't point to when it happened.
The pace the last few years definitely seems rapid, just don't want there to be a false impression.
There is still plenty of improvement ahead but I don’t think anything genuinely surprising will come from the current regime of feedforward models. What is missing is action — an interactive feedback loop between agent and environment. Progress in RL and robotics has been very slow by comparison and unless we see a breakthrough there, I would guess the GPT phase plateaus in the next 5-10 years.
Haven’t we pretty much figured out these things are doing more than just predicting the next token at this point?
There’s probably a lot to be done with a “prediction machine”, birds aren’t all that smart but can catch bugs in midair.
It’s actually a good question: what will next-token prediction be able to do on new datasets? The error is thinking you can answer it, even in broad terms.
For example, I expect skill at writing some kinds of code to improve dramatically because running tests in a sandbox looks easy. It’s already being researched. [1] Extending that to device drivers might be a bit harder. Fuzzing is already mostly automated and smarter fuzzing could get pretty scary.
[1] https://nanothoughts.substack.com/p/reflecting-on-reflexion
If you follow the link I shared, some researchers automated asking GPT4 to write tests, running the tests in a sandbox, and feeding the results back in.
With access to python interpreter I don't see any issue
Transformers are Sample-Efficient World Models:
https://github.com/eloialonso/iris#transformers-are-sample-e...