Not an expert in this space.
Aren't tokens transformed with position dependent information in most models?
I believe llama applies a rotation to the vector based on the position in the input.
Aren't tokens transformed with position dependent information in most models?
I believe llama applies a rotation to the vector based on the position in the input.
https://github.com/huggingface/transformers/blob/222505c7e4d...