On a related note, one thing I still do not understand is why are positional encodings 'added' to the token embeddings as opposed to (having a smaller position encoding vector that is) 'concatenated'. It would be great if someone could explain.
h_i = attention(W_i^Q Q^T @ W_i^K K) W_i^v V
h = W_o @ concat(h_1...h_8)
Seems to me like it’d be a quite low number compared to the dimensionality of the semantic vectors?