DroPE: Extending the Context of LLMs by Dropping Their Positional Embeddings
pub.sakana.ai
pub.sakana.ai
I wish this point was explained further instead of being just a footnote. It seems like the central insight that is essential for this technique to work, and it is not obvious to me, maybe because I haven't implemented a transformer from scratch.