The Impact of Positional Encoding on Length Generalization in Transformers
arxiv.org
arxiv.org
Interestingly, position embeddings seem unnecessary for decoder only architectures going against what is conventionally thought. Wondering, if this result is as strong as it sounds.
> Overall, our work suggests that explicit position embeddings are not essential for decoder-only Transformers to generalize well to longer sequences.