Once you realize this "Multi-headed Attention" is just kernel smoothing with more kernels and doing some linear transformation on the results of these (in practice: average or add)!
0. http://bactra.org/notebooks/nn-attention-and-transformers.ht...
Once you realize this "Multi-headed Attention" is just kernel smoothing with more kernels and doing some linear transformation on the results of these (in practice: average or add)!
0. http://bactra.org/notebooks/nn-attention-and-transformers.ht...
> To resolve these issues, we introduce the Performer, a Transformer architecture with attention mechanisms that scale linearly, thus enabling faster training while allowing the model to process longer lengths, as required for certain image datasets such as ImageNet64 and text datasets such as PG-19. The Performer uses an efficient (linear) generalized attention framework, which allows a broad class of attention mechanisms based on different similarity measures (kernels). The framework is implemented by our novel Fast Attention Via Positive Orthogonal Random Features (FAVOR+) algorithm, which provides scalable low-variance and unbiased estimation of attention mechanisms that can be expressed by random feature map decompositions (in particular, regular softmax-attention). We obtain strong accuracy guarantees for this method while preserving linear space and time complexity, which can also be applied to standalone softmax operations.
1: "Context Is The Next Frontier by Jacob Buckman, CEO of Manifest AI" (https://youtu.be/wJyl6kBCwmY?si=ruxMWdENjazu3rp6)
∑ᵢ yᵢ · K(xᵢ, xₒ) ⁄ (∑ⱼ K(xⱼ, xₒ))
In regular attention, we let K(xᵢ, xₒ) = exp(<xᵢ, xₒ>).Note that in Attention we use K(qᵢ, kₒ) where the q (query) and k (key) vectors are not the same.
Unless you define K(xᵢ, xₒ) = exp(<W_q xᵢ, W_k xₒ>) as you do in self-attention.
There are also some attention mechanisms that don't use the normalization term, (∑ⱼ K(xⱼ, xₒ)), but most do.
That clarifies things...
> 0. http://bactra.org/notebooks/nn-attention-and-transformers.ht...
Reductions of one architecture to another are usually more enlightening from a theoretical perspective than a practical one.
Now you got the definition, and the alternative naming you can use to search for resources.
1: "Context Is The Next Frontier by Jacob Buckman, CEO of Manifest AI" (https://youtu.be/wJyl6kBCwmY?si=ruxMWdENjazu3rp6)