- how are the input encodings generated?
- what is in those position vectors?
- how are the attention vectors learned?
The answer is that these things are all learned as the network is trained; the whole thing is one “thing”. The concept that is most important in understanding neural networks generally is that they start out as just a bunch of random numbers and then the numbers are gradually adjusted until the outputs converge closely enough on the desired loss.
I recommend watching Karpathy’s YouTube video where he codes up a Transformer from scratch. It’s the best way to understand these beasts.