Unlimiformer: Long-Range Transformers with Unlimited Length Input
arxiv.org
arxiv.org
2. This idea is quite similar to retrieval transformers and Hopfield networks which have been known and published for several years now. It's not really that novel.
3. Due to the preceding points, the title can easily mislead people. It's not really a conventional transformer, and it's not a breakthrough.
4. This paper is a preprint and not peer-reviewed.
"I generally don't enjoy seeing preprints like this going to the top of Hacker News. This would be a higher quality submission if the paper was peer-reviewed or put into a greater context, like a blog post discussion or something like that."
Let me retract this and say something a bit nicer :) I personally think there this specific preprint making it to the top of HN is potentially harmful, because of the hype around LLMs, the diverse audience of readers here, and the specific title that implies a claim of "transformer with unlimited context length", when this is misleading. I don't have anything against preprints in general - a lot of work outside of the peer-review process ends up being very impactful.
Is it? I had thought retrieval transformers "merely" used retrieval as a backend of sorts rather than a substitute for the attention itself?
[0] https://jalammar.github.io/illustrated-retrieval-transformer...
Your comment essentially says - this is not a high quality submission because readers might not actually read it, which is no fault of the work, or submitter.
I'd argue that on average, most readers won't have a good enough understanding, or read the paper far enough, to understand that the reality is closer to "it's not a breakthrough" rather than "Transformers with Unlimited Length Input".
So, I wholeheartedly welcome this type of hype-breaking leading comment.
The headline will always be "BATTERY BREAKTHROUGH PROMISES TO ROCKET ELON MUSK TESLA TO THE MOON!!!" and while it's easy to know that some amount of cold water is necessary you need to spend a nontrivial amount of attention and have a nontrivial amount of knowledge to figure out just how much cold water. It's a useful thing to outsource. Did a research group see outperformance in an experiment with 1% probability of translating into production? Or is CATL scaling up a production process? The "well actually" comment will contextualize for you. If there's a "well actually" reply to the "well actually" comment, that tells you something too. Upvotes/downvotes dial in the distributed consensus.
It's far from perfect, but I'd challenge detractors to point to a more effective method for large-scale democratic truth seeking.
it seems some technical and social countermeasures could be deployed so there was at least token visiting of a link before commenting was allowed, which court raise discourse in this forum, at least. It's a false dichotomy to only consider the two extremes of Peer Reviewed and HN reviewed. In particular, as mentioned up thread, the incentive to do peer review (or replicate an experiment, for that matter) isn't as high as working on your own research, attempting to do some novel work for a shot at a Nobel prize.
as every coder who's been the reviewer on a code review knows, it's difficult, and one that often is not well prioritized among other priorities a senior IC might have, leading to poor quality reviewing or other work slipping (only so many hours in a week). Thus, one could imagine a system that gives direct compensation for review work to grad and smart undergrad students who work in the area being discussed vetting and/or contextualizing claims like "this work is/is not novel", rather than hoping someone who does work in that area is procrastinating on HN at just the right time to make the claim and the rebuttal and the rebuttal-rebuttal. If the rubuttal is posted after the thread falls off the front page, is anyone not in that thread even going to know that what they read was wrong?
That's how I feel, anyway. I'd rather have seen a comment that has the same explanations in it but just generally less grumpy! Saying stuff like "It's not really that novel." doesn't really contribute much, when it could either be explained why it isn't novel by explaining how similar it is to something earlier that can be referenced, or thinking about what if anything is novel in this research - assuming it isn't being accused of just replicating something already done.
But yes, you're correct in this instance that it's not necessarily 'huge news' since it is highly similar to a long list of the Reformer (LSH-based), Performer (FAVOR**), FNet (Fourier-based), Routing Transformer, Sparse Transformer, Longformer (task specific sparse), blockbert, XLNet/xfmr-xl (slide + relative PE), BP-Transformer (binary partition), BigBird (global and rand attn), RWKV which is..., etc.
** FAVOR actually is innovative and different in this space, but towards similar ends anyway
Anyone who has been for whatever reason reading the papers since 2017 has invariably read dozens of these papers.
Anyone who has heard of GPT-x in 202x and started from there probably didn't.
This will likely change with implementation of memory retrieval, some form of linear attention etc. in many productions models, and the democratization of some decoder models... although I have been thinking this for a while.
Don't get me wrong, you want to hire the people who know these papers, especially if they started after 2017 :-)
When something works well, I don't care much about point 4.
Personally, I've had only mixed success with KNN search on long sequences. Maybe I haven't done it right? I don't know. In my experience, nothing seems to work quite as well as explicit token-token interactions by some form of attention, which as we all know is too costly for long sequences (O(n²)). Lately I've been playing with https://github.com/hazyresearch/safari , which uses a lot less compute and seems promising, though it reminds me of things like FNet. Otherwise, for long sequences I've yet to find something better than https://github.com/HazyResearch/flash-attention for n×n interactions and https://github.com/glassroom/heinsen_routing for n×m interactions. If anyone has other suggestions, I'd love to hear about them.
This opinion seems totally backwards to me. I'm not sure what you think peer-reviewed means? Also I prefer full preprints than blog posts. But then again, I have no idea why ones like the daily blogposts of Seth Godin (to pick on one randomly, sorry it's not personal) so often go to the top of hacker news. Maybe opinions like yours explains it?
I agree.
> I'm not sure what you think peer-reviewed means?
Posting to HN is a form of peer-review, typically far better than the form of "peer-review" coopted by journal publishers.
This is a rather self-aggrandizing view, and I think it speaks to the level of ego that underpins a lot of the discussion on here.
However, if you limit the question to "papers which make bold conclusions of the type that generates lots of discussion on HN", I think HN will be more likely to find methodological flaws in those papers than the peer review process would. I think that's mostly because papers are optimized pretty hard not to have any problems which would cause them to be rejected by the peer review process, but not optimized very hard to not have other problems.
Which means, on average, I expect the HN comment section to have more interesting feedback about a paper, given that it's the sort of paper that gets lots of HN discussion, and also given that the author put a lot of effort into anticipating and avoiding the concerns that would come up in the peer review process.
Which, to a reader of HN, looks like "a lot of peer-reviewed papers have obvious flaws that are pointed out by the HN commentariat".
I do think, on the object level, a pre-print which the author intends to publish in a reputable journal will be improved more by fixing any problems pointed out by HN commenters than by fixing any problems pointed out by peer reviewers, and as such I think "post the pre-print on HN and collect feedback before peer review" is still a good step if the goal is to publish the best paper possible.
Although I should add I have no background in academia and don't feel prepared to have that discussion.
It'll cause some journals to not publish your work.
Also, I want to remind everyone that ML uses conferences as the main publishing mechanism, not journals. While things like JMLR exist, that's not where papers are targeting.
Maybe we just need to let researchers evaluate works based on their merits and not concern ourselves with things like popularity, prestige, and armchair experts' opinions. The latter seems antiscientific to me. We need to recognize that the system is noisy and Goodhart shows us we aren't optimizing merit.
[0] an example is that I had a strong reject with 2 lines of text. One stating that it wasn't novel (no further justification) and the other noting a broken citation link to the appendix. No comments about actual content.
[1] As another example, I've had reviewers all complain because I didn't compare one class of model to another and wasn't beating their performance. I beat the performance of my peers, but different models do different things. Image quality is only one metric. You wouldn't compare PixelCNN to StyleGAN.
Ok, but how would the researchers communicate their evaluation to non-experts? (Or other experts who didn't have the time to validate the paper)
Isn't that exactly what a review is?
My impression is the armchair experts are more likely to be found on HN.
Conferences, journals, and papers are not for non-experts. They are explicitly for experts to communicate with experts. The truth is that papers have never been validated and likely never will. Code often isn't uploaded alongside papers and when it is I know only a handful of people that look at it (including myself) and only one that executes it (and not often). Validation only happens with reproduction (i.e. grad students learning) and funding doesn't encourage that. Despite open source code, lots of ML is still difficult to reproduce, if it can be done at all.
We also use normal forms of communication like Twitter, HN, Reddit, email, etc but there's a lot of noise (as you note). We speak a different language though, so you can often tell.
Frankly, a lot of us are not concerned with explaining our work to layman. It's a lot of work, especially the more complex a subject is and we're already under high pressure to continue researching. It's never good enough. There's no clear "done with work" time in jobs like this. You're always working, so allocate your energy (I'm venting and mentally fatigued right now). I used to be passionate about teaching laymen but I'm tired of arguing with armchair experts. Still happy and passionate about teaching my students and performing research, so that's where I'll spend most of my energy: in the classroom or blogs. The more popular a subject is, the more likely this is to happen too, ironically.
Communication should come from news, university departments, and specialty science communicators, but that's broken down. Honestly, I just think it's a tough time for laymen to get accurate information. There's a lot of good information out there for you all (us researchers learn from publicly available materials) but expertise is being able to distinguish signal from noise, and the greater the popularity, the greater the noise. This isn't just true for ML, we see this in things like climate, nuclear, covid, gender/sexuality, and other hot topics. Only thing you can do is actually use a common strategy from researchers: have high doubt and look for patterns from research groups.
If you go to a computer science conference you might talk about the headliners later but you actually learn a lot from talking to less famous people at the back of the room, scanning large numbers of poster papers, sharing a bottle of wine at dinner with four people and having one of them get way too drunk and talk trash about academics you distantly know, etc.
Lower-quality papers on arXiv give me a bit of that feel.
I would say the primary difference between a conference peer review board and HN is that the author is obliged to respond to the reviewers on the board. I would not say there’s any particular difference in qualifications.
That already narrows it down greatly compared to the general public you find on the internet.
I'm not so sure about that. I've read a lot of things that should have never left peer review or editing stages, while some of the most impotent papers for my field never left preprint.
Overall I think the most imprortant step of peer review is you as the reader in the field. Peer review should catch the worst offenders out, saving us all some time, but it should never be viewed as a seal of approval. Everything you read should be critically evaluated as if it were a preprint anyway.
I don't think "participate" and "leave a comment" are the same thing. A random person most likely wouldn't be able to follow or contribute to the conversation. They could only leave a comment.
It's a bit pedantic, but noise usually sinks to the bottom.
I mean, hypothetically, this whole thread could be stuffed with sock puppet accounts of the author. How would you know?
Except of course for Nature-level papers, but most people never get to review papers like that
Novelty is especially a joke. ViTs are "just" NLP encoding transformers. T2I models are "just" NLP models connected to generative models. Diffusion models are "just" whitening models. GPT3 is just GPT2 with more layers and more data which is just GPT with more layers and more data. We can go even deeper if we pull from math and physics works. But that doesn't mean these works haven't been highly fruitful and useful. I'm happy all of these have been published.
> because of the hype around LLMs
I too hate the hype, but it is often bimodal. There are people who are far too critical and people who are far too accepting. The harm is not preprints or people reading papers, the harm is people who have no business/qualifications evaluating works confidently spouting out critiques. It is people not understanding that researchers are just critical of one another's work by default and that doesn't mean it shouldn't have been published.
It is well known that reviewers are good at identifying bad papers but not good at identifying good papers[0,1]. Which let's be honest, that means reviewers just have high reject rates in a noisy system. Making publication as a metric for merit a highly noisy one at best.
As for the paper:
Many LLMs and large models are using attention approximations. Nor is the kNN technique particularly new. My main complaints are a lack of comparisons for Figure 3 and 4, but I'm not a NLP person so I don't even know if there's some other good works that can compare better (BART is a common baseline). But generative models are (unfortunately not notoriously known) extremely difficult to evaluate. Paper seems fine to me. It is useful to the community. I don't like the name either, but their input is limited by computer memory, not the model. I would want to see more on this. Not a NLP person all I can say is that this looks neither like a strong reject nor a strong accept. I'll leave it to the community to determine if they want more experiments for the conference publication or not but the work seems useful.
[0] https://inverseprobability.com/talks/notes/the-neurips-exper...
In the transformer architecture one has to compute QKT.
QKT=(hd * Wq * WkT)heT (equation (2) page 3 in the paper).
Where hd is the hidden state of the decoder, and he is the hidden state of the encoder, and Wq and Wd are some parameters matrices, and T denotes the transposition operation.
By grouping the calculation this way, in a transformer encoder-decoder architecture, they can build and use only a single index (you index the he vectors using a vector database) for all the decoder layers queries. Instead of having to build 2 * L * H indices (with L the number of layers of the decoder and H the number of head in the decoder).
But what makes it a little dubious, is that this transformation mean you make your near neighbor queries in a space of dimension "dimension of the hidden state", instead of "dimension of a head" that is H times smaller.
So if you had to build 2 * L * H indices each index would be H times smaller.
So you only gain a factor 2 * L. But the trade-off is that you are doing a near neighbor search in higher dimension where you are then subjected to the curse of dimensionality (the higher the dimension the more similar all points are to each other). Whereas the whole point of projections in transformer is to lower the dimension so that the knn search make more sense. So to get the same accuracy, your near-neighbor search engine will have to work a lot harder.
Also as an approximation of the transformer, because it's using some knn search, it comes with the problems associated with it (for example harder to train because more sparse, and a tendency to hyperfocus), but it can be complemented with low-rank linearization of the attention to also have the neural net act on the gist rather than the closest neighbors.
> So you only gain a factor 2 * L. But the trade-off is that you are doing a near neighbor search in higher dimension where you are then subjected to the curse of dimensionality (the higher the dimension the more similar all points are to each other).
I thought that the curse of dimensionality meant that in higher dimension, points got farther apart
Sounds like you are both right?
You only need the important concepts, not individual words
If I'm understanding the article, this approach would not use the skipped words to influence the output, so I thinks ita necessarily different.
Point (a) is extremely hard to discern, especially when people are chasing third-significant-digit gains on common benchmarks; it's essentially multiple-testing false discovery in action. I've seen whole families of methods fail to transfer to new domains...
Point (b) is also a real issue. As you increase the number of bells and whistles, each with their own hyperparameters with non-linear impacts on model quality, it becomes impossible to say what's working or not.
In practice, i think we see some cycles of baroque incremental improvements, followed by someone spending a year stripping away the bullshit and getting something simple that outperforms the pack, essentially because it's easier to do hyperparam search over simpler models once you figure out the bits that actually matter.
The Unlimiformer paper is about a new way to make computer programs that can summarize really long pieces of text. Normally, when you ask a computer program to summarize something, it can only handle a certain amount of text at once. But with Unlimiformer, the program can handle as much text as you want!
The way Unlimiformer works is by using a special technique called a "k-nearest-neighbor index" to help the program pay attention to the most important parts of the text. This makes it possible for the program to summarize even really long documents without losing important information.
Overall, Unlimiformer is an exciting new development in natural language processing that could make it easier for computers to understand and summarize large amounts of text.
Personally I think the RNN/LSTM state handling approach is going to be something we revisit when trying to advance past transformers. It handles state in a way that generalizes and scales better (it should in theory learn an attention-like mechanism anyway, and state is independent of input size).
It may be harder to train, and require further improvements, but it really seems more like an engineering or cost problem than a theoretical one. But I’m only an amateur and not an expert. Maybe continued improvement on attention will approach generalized state handling in a way that efficiently trains better than improvements on more generalized stateful approaches improve training.
Because of exactly that.
Also the attention mechanism is baked in during pretraining. So whatever max context length you want increases the compute cost of training by at least a function of said "bad complexity." Even just 4096 tokens of max context is much more expensive to train than 2048. So if we want models with 8k, 32k, or more context then the training costs get out of hand quickly.
IIUC, this is no longer necessarily true with positional encodings like ALiBi: https://github.com/ofirpress/attention_with_linear_biases
It seems mostly like a vertically integrated vector DB + existing LLM call, but correct me if I'm wrong. There are of course some performance gains with that, but the holy grail of "understanding" at unlimited length still seems unsolved.
Most vector DBs use (at least) some kind of KNN anyways.
Linking to a different submission of the same link with 0 comments doesn't add anything.
But unlike sites like Reddit, with the exception of self / ask HN / etc posts, nobody really pays attention to who the submitter is, so enjoy the conversation finally breaking out on it as consolation for not getting karma points, but skip linking to dead submissions :)
FYI, if you ever submit something that fails to get any traction / upvotes, then I've seen mods say before (@dang will hopefully correct me if I'm wrong) that a) it's OK to try submitting a second time maybe after a day or so (but not keep submitting over and over) or b) send the mods an email with a brief reason why it's a link that should interest HN readers for it to be potentially added to a "second chance pool". Though in the case of this link, between three of you it was posted two days ago, one day ago, and today which has finally got a bit more notice, so worked out alright in the end :)
The length of inputs is theoretically bounded by the memory limitations of the computer used. More practically, using a CPU datastore is many times slower than a GPU datastore because of slower search and the need to transfer retrieved embed- dings to the GPU... (continues)
1) Normalize input (batch norm, 2015)
2) Competitive dynamics / lateral inhibition (softmax in attention layers, 2017)
3) Cluster best matching activation vectors (top-k keys, 2023)
https://arxiv.org/pdf/2305.01625.pdf
> Unlimiformer summary:
> The first part of the novel focuses on the question of whether or not the Russian nobleman, Dmitri Fyodorovitch, has killed his father. In the town of Ivanovna, the lieutenant-colonel of the Mushenkhanovitch is accused of the murder of his brother Ivanovitch. The lieutenant-incommand, Vasilyevitch, takes the form of a dog, and the two men–the two men and the woman who are questioned by the court-martial–murphy. The two men cry out to the God of Russia for help in their quest to save the town. The man, afraid of the wrath of the God, hands the dog a bunch of letters that are supposed to be proof of his love for his brother. The old man–the one who had killed his mother, and then found the letter–arrives. He reads it–asked the old man to forgive him for the murder and then takes the dog away. The other men, all of whom are prisoners, demand that the man confess his crime to the court. The first and most important thing they tell the court is that they love the man. The court acquits the man and sentences the man to death. The second man–an old officer of the town, Alekandrovitch–askes to tell them the same thing. The third man–in the process of confessing his crime–is Vashenka, a drunk man who has been sent to the town to kill his father, for reasons which are not entirely clear to the people. The woman’s servant, Evgenyevna, is also the one who has told the court the story of the Medvedevitch’s murder, for the good old man’s and the young man’s love. The three men, who are separated for the first time, are laughing at the man’s attempt to seduce Mitya. The young man, in the meantime, is conscripted into the town-side. He tells the court that he loves her, but he has yet to tell her the true story. The men, in this room, demand a man to kill her, and she will not betray them. The women, in their own country, are rebelling against the man who had sent them three thousand roubles, and they will not allow the man of the people to see them. They will not let the man in the town be allowed to see the man–or Dmitriovitch; he will have her husband killed him. He will not tell the people who love him. The next man, named Vashenovitch, arrives, and takes the man away. They all begin to laugh at the fact that he has succeeded in seducing and entrusting his brother Dmitri. He is then taken away to the old woman’s house, where the governor-side-of-the-world, and his sister, Arkadin, is being punished. The priestesses and the baron are shocked, for they have been so virtuous and well-suited. The only thing they will be able to do is kill the priest. They threaten to burn the priestess to death, for she has been so wicked and libidinous that she has not yet seen the priest, for her husband. The priests–ostensibly convinced that she is a woman who loves the priest and has been punished for her love and for allowing the priest to marry her. The last man, Yakivitch, arrives at the house, and, after a long day of drinking and then some of the men–is killed. He and the priest are ordered to leave the town so that the priest can finally be reunited with the people of the old lady. The final man, the commander of the St. Petersburg town of Arkadina, is sentenced to death for the crime of having killed and then the lieutenant of the governor, for taking the money. The commander, the former lieutenant-delegation of the People’s Army, is summarily executed, and all the men, except for the commander, have been summarily punished for their crime. The entire town is shocked and, in a very dramatic way, the priestesses plead for the forgiveness of the man, for allowing them to kill and imprison Ivan. They plead for their brother to be restored as well, for all the people they have loved, and for the priestor to tell the story
I'm in the process of implementing a framework based on this idea.
I have written a paper on this recently, https://arxiv.org/abs/2302.01834
I have a discord channel https://discord.cofunctional.ai.