As you read each panel of a comic book, you don't just look at the words in the speech bubbles, but you also pay attention to who's talking, what they're doing, and what happened in the previous panels. You might pay more attention to some parts than others. This is sort of like what the Transformer model does with text.
When the Transformer reads a sentence, it doesn't just look at one word at a time. It looks at all the words at once, and figures out which ones are most important to understand each other. This is called "attention." For example, in the sentence "The cat, which is black, sat on the mat," the Transformer model would understand that "cat" is connected to "black" and "sat on the mat."
The "attention" part is very helpful because, like in a comic book, understanding one part of a sentence often depends on understanding other parts. This makes the Transformer model really good at understanding and generating language.
Also, because the Transformer pays attention to all parts of the sentence at the same time, it can be faster than other models that read one word at a time. This is like being able to read a whole page of your comic book at once, instead of having to read each panel one by one.