Transformers from Scratch (2019)
peterbloem.nl
peterbloem.nl
> In the context of neural networks, attention is a technique that mimics cognitive attention. The effect enhances the important parts of the input data and fades out the rest—the thought being that the network should devote more computing power to that small but important part of the data. Which part of the data is more important than others depends on the context and is learned through training data by gradient descent.
Now, it remains to figure out what the "self" if "self-attention" is for, but maybe someone else can fill that gap. ML isn't something I care about particularly, so don't want to spend too much time.
....but this is not that.