Wasn't there a paper yesterday that turned context evaluation linear (instead of quadratic) and made effectively unlimited context windows possible? Between that and 1.58b quantization I feel like we're overdue for an LLM revolution.
https://arxiv.org/html/2404.08801v1 Meta Megalodon
https://arxiv.org/html/2404.07143v1 Google Infini-Attention
https://arxiv.org/html/2402.13753v1 LongRoPE
and a ton more