While Transformer's Self Attention(SA) is great, there are many applications where SA doesn't apply. For a more comprehensive overview of attention mechanisms, I often find myself coming back to Lilian Weng's post:
https://lilianweng.github.io/lil-log/2018/06/24/attention-at...