This is the intuition that worked for me: transformers are mini search engines!
Each attention node is essentially performing a simple search on the previous text given a single word. It's kind of like Google! The actual search or "query" is the word you're looking at, the page title or "key" is the other words in the sentence, and the page/document or "value" is... also the other words in the text (each word is like its own page title). So if you imagine the whole internet is just the sentence "Hi my name is Bob", and you googled "name", you should get a results page with "Bob" at the top, followed by "my" then "is" then "Hi".
As an optimization, the transformer will actually run this "search" operation for all the words in the sentence at the same time. It then spits out a matrix saying how relevant each word is too every other word. Basically each transformer in your neural network is like a little guy who says "Hey, if you're looking at X word, here's the other words Y and Z that might be relevant in the same sentence". The rest of the NN can then use that information using more typical feed forward techniques (I'm mostly glossing over use of embeddings and other various special techniques that are used in practice here). There will be multiple transformers in a given LLM that each perform a slightly different search.
But how does a transformer know how to rank these words? In a nutshell, as we train it to predict text, signals from the error in that training prediction back propagate to the transformer node, and it gradient descends to be the best search engine it can be. Each transformer will have a slightly different random starting state and different distribution of modules further forward in the network back propagating error to it, and thus they will each converge to a slightly different "search algorithm" so to speak.