Tensor Product Attention Is All You Need
arxiv.org
arxiv.org
When trying to deploy llms in with larger context windows constrained environments 2 things start to hurt: a) increased memory footprint for longer KV cache b) increased decode speed due to longer context window. this paper addresses a) only, which is useful, but we are still left with b) (right?)
> These variants illustrate TPA’s versatility in balancing memory cost, computational overhead, and representation power. By choosing which dimensions (heads or tokens) remain contextual and adjusting ranks (RQ, RK, RV ), TPA unifies multiple existing attention mechanisms— such as MHA, MQA, and GQA—under one framework, while potentially reducing the KV cache size by an order of magnitude during autoregressive inference.
re: the title, it might be the true one if their proofs hold up
---
I'm now curious if the Element-wise Attention is All You Need preprint can be fit into this framework. Sadly my math is not currently up to the task. It appears to offer even better computational savings during both training and inference while maintaining accuracy, though only tested with a smaller model
I haven't read this paper (yet) but isn't this the case that mostly applies to training and not so much to inference? A good example would be flash-attention, it trades the higher flops for better memory utilization but it's mostly irrelevant in inference workloads.
They also, interestingly, don't compare against the flash-attention. Flash-attention outperforms all of the other attention mechanisms mentioned in the paper: MHA, MQA, GQA, and MLA.
Both datacenter grade CPUs and GPUs have enough compute to carry out the self-attention computation but it is only the latter that has enough hi-bandwidth memory to make the algorithm really perform. If this hadn't been the case, the theory behind flash-attention wouldn't materialize, and it does, and reason being that (main) memory is slow.
Deep FFWD networks OTOH are compute-bound.
For N self-attention layers, there will be N compute (tensor) units doing the computation in parallel. To retire the computation, each compute unit will need to LOAD/STORE from and to the chip memory. At batch size B, this only becomes a bigger scale, e.g. B * (N, LOAD/STORE).
But assuming your KV cache size is << model size, that simplification is pretty accurate.
See, e.g. https://www.databricks.com/blog/llm-inference-performance-en...
You can just scroll to the first chart they have that explains the idea.
This is kind of a theme in HN now. The top comments are completely besides the point of the article/story/etc.
Tensor Product Attention: A Memory-Efficient Solution for Longer Input Sequences in Language Models
Maybe a Greasemonkey script to pass arXiv abstracts to a local Ollama could be something...
*side effects TBD
https://scholar.google.com/scholar?hl=ro&as_sdt=0%2C5&q=%22i...
Always has been.
silent gunshot
see Section 3.4
But, on the other hand, it's hard to get researchers to read your paper, esp. in fast-moving areas. Every little thing might be the difference between reading the abstract or not. Reading the abstract might lead to reading the intro. And so on.
So, for better or worse, the competition for human eyeballs is real.
Ironically, in this case, "attention" is all that the authors want.
Bloody hell and brimstone. Been crazy 57 years and a half already.
There's another paper I saw yesterday, "Element-wise Attention is All You Need" which looks like an early preprint, written by a solo author with a solo A800, and tested on some smaller problems. If the results hold up for language benchmarks, it could reduce resource requirements during training as well. It looks to have a lower complexity when scaling
There is something in the github repo about higher-order decompositions. Don't find where the method for factoring is given.
> Specifically, for each token t, with a small abuse of notation, we define:
Trading computational complexity for space.
"The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely."
They did not invent attention, but while previous language models had used attention as an auxiliary mechanism, they removed everything but the attention and the models still worked. Really, the title already says it all.
In theoretical CS you have state machines, pushdown automatons and turing machines. What may surprise you is that the difference between those three does not lie in the way the algorithm for them is represented. They all use state transition diagrams!
Pushdown automatons are more powerful than state machines, because they have a stack onto which they can push new data, peek at the top of the stack or pop data off the stack.
Now here is the kicker! How do you get a turing machine? You take a pushdown automaton and remove the restriction that you can only push, peek or pop at the head! You can now move the head pointing at the stack, your stack has turned into a tape!
The key difference lies in the datastructure that the transition diagram is manipulating, not the algorithm itself!
The attention mechanism of the transformer architecture is analogous to a tape that can only be written to once per blank and in theory that alone is enough to emulate a full Turing machine with read, write, rewrite and delete semantics.
Why do every paper has to mention this word "novel" and these titles are getting crazier day by day.
I thought the number of parameters grows quadratically with context window length - what do they mean?
No.
I hate ads, but I’m not paying for YouTube Premium either. That’s how it goes. I get ads.
No that's not how it goes. You get Ublock Origin and then you don't get ads. Simple as that.
If you don't like ads and don't fight against them it means you accept ads and want to see more of them shoved down our collective throat. At least from the perspective of marketers and industries who rely on ads. That's how we ended up in this predicament in the first place. Lazy compliance.
If Youtube isn't sustainable without ads, it should die so that natural selection can take over and so that a better ecosystem can finally take its place. Every single "Youtuber" hates the platform, mostly because it has zero transparency being a Google product. The viewers hate it too because it constantly takes down their favorite videos and creators and because it's full of ads.
The only reason it's (still) the main site for hosting videos is quite literally just ad-fueled inertia due to the intrinsic cost of hosting videos. If ads didn't exist the only sustainable solution would be something less centralized like Peertube. And to me that's a desirable outcome.
This doesn't apply to arXiv though, as it is not peer reviewed nor edited and is funded by various institutions.
Admittedly the academic publishing system is so corrupt that it's hard to phantom, so it's easy to misunderstand it.
Typically you do pay for the papers (and publisher profits), either through taxes or inflated product prices.