I wonder if these N-gram reduced models, augmented with confidence measures, can act as a very fast speculative decoder. Or maybe the sheer number of explicit rules unfolded from the compressed latent representation will make it impractical.
I think there will also be benefits to that both in interpretability and hardware acceleration. In time, maybe cheaper pretraining of useful models.
[0] https://transformer-circuits.pub/2021/framework/index.html
1: https://docs.vllm.ai/en/latest/features/spec_decode.html#spe...
This, as I found out from this repo [0] linked in the Twitter thread in the documentation (which for some reason they didn't just link to directly), seems to be a regular Markov chain of context, if it even builds a stochastic matrix. See algorithm below.
Current prompt
"Article: (CNN)French striker Bafetimbi Gomis, who has a history of [...]
Summary: French stri"
Prompt lookup algorithm
1. Get last few tokens from prompt -"French stri"
2. Search for "French stri" in prompt
3. Match found - return next k tokens after match as candidate completion -"ker Bafetimbi Gomis, who has"
Candidate tokens
"ker Bafetimbi Gomis, who has"
[0] https://github.com/apoorvumang/prompt-lookup-decoding