I think your misunderstanding is that fully connected layers can operate in the same way that attention does — they can’t. A fully connected layer operates on one dimension at a time. Typically language models have two (plus a batch dimension). One dimension is your token/word dimension. The second is the hidden dimension.
The hidden dimension is constructed inside the model to create space where it can embed each token into a vector and then enrich that space with contextual information derived from the sequence. In order for that to occur, the model must have a means of transferring information from along the token dimension.
One way to accomplish this is to use a 2d convolution; however, the scope of a convolution is limited to the size of its kernel. A fully connected layer is the same as a 2d convolution with a kernel size of 1. So you can see that no information from neighboring tokens can be applied to the hidden space.
The standard self-attention equation has a global scope from the full matrix multiplication of the input tensor with its transpose. Each element of the resulting matrix demonstrates some interaction with every other token in the sequence. Next a softmax operation is applied, which acts as a gating or relevance function. Finally, this is multiplied back to the original input to build that information into the hidden dimension for each token.
There have been attempts to do similar operations using fully connected layers. Look at the architecture of SGUs (spatial getting units). In some applications, they have good performance, but because fully connected layers operate on each dimension independently and serially, they are not equivalent to attention.
Last, my best recommendation for anybody trying to understand attention is to stop reading articles and instead spend your time looking at the math. It’s usually much less confusing than any of the dozens of explanations floating around the web, including the one I just gave. The math is not too complicated, especially once you know the reasons for why we need to use it.