The Transformer is unique because it uses a mechanism called "attention" to understand the relationships between words in a sentence, which works like this:
(1) Encoding: First, the Transformer turns each word in a sentence into a list of numbers, called a vector. These vectors capture information about the word's meaning.
(2) Self-Attention: Next, for each word, the Transformer calculates a score for every other word in the sentence. These scores determine how much each word should contribute to the understanding of the current word. This is the "attention" part. For example, in the sentence "The cat, which is black, sat on the mat," the words "cat" and "black" would get high scores when trying to understand the word "black" because they are closely related.
(3) Aggregation: The Transformer then combines the vectors of all the words, weighted by their attention scores, to create a new vector for each word. This new vector captures both the meaning of the word itself and the context provided by the other words in the sentence.
(4) Decoding: Finally, in a task like translation, the Transformer uses the vectors from the encoding phase to generate a sentence in the target language. It again uses attention to decide which words in the original sentence are most relevant for each word it's trying to generate in the new sentence.
One key advantage of the Transformer is that it can calculate the attention scores for all pairs of words at the same time, rather than one at a time like previous models. This allows it to process sentences more quickly, which is important for large tasks like translating a whole book.