On the other hand:
> There are various forms of attention / self-attention, Transformer (Vaswani et al., 2017) relies on the scaled dot-product attention: given a query matrix , a key matrix and a value matrix , the output is a weighted sum of the value vectors, where the weight assigned to each value slot is determined by the dot-product of the query with the corresponding key
There HAS to be a better way of communicating this stuff. I'm honestly not even sure where to start decoding and explaining that paragraph.
We really need someone with the explanatory skills of https://jvns.ca/ to start helping people understand this space.