SepLLM: Accelerate LLMs by Compressing One Segment into One Separator
sepllm.github.io
sepllm.github.io
But if it is true that the separators contribute the most towards the attention scores, wouldn't that imply that the tokenization scheme can be improved? Introducing a compression scheme seems like patching around that compared to if the model naturally generated a more random attention distribution.
'Why waste time say lot token when few token do trick?"
-Kevin Malone