HyperAttention: Long-Context Attention in Near-Linear Time
arxiv.org
arxiv.org
"when half of all attention layers are patched (i.e., 14 layers), we verify that most of the tasks do not degrade more than 13%."
According to the paper, for most tasks it reduces benchmark scores substantially. Perhaps to the point where a smaller model would yield better inference time and higher benchmarks.
However, summarization benchmarks see almost no degredation, great!
And I think the optimized backends should implement that sliding 16k context soon...
Anyway, point is a huge context really helps certain types of queries, and VRAM usage is reasonable with a 7B model.
They tried something, it improved some things and made other things worse.
Also publish yet another "SOTA framework" thats barebones and won't be maintained for very long!
The researchers who published the negative CFG LLM paper made an earnest effort to pull it into the popular frameworks. That really stands out in my memory, that is incredibly rare.