Lossless Acceleration of LLM via Adaptive N-Gram Parallel Decoding
arxiv.org
arxiv.org
I do have a feeling of dejavú, like I've seen this before on hn.
You're either thinking of speculative encoding more generally, or Medusa: https://arxiv.org/abs/2401.10774
From the below Modal link they use: >continuous batching, so multiple generations can take place at the same time on a single container
>PagedAttention, which applies memory paging to the attention mechanism’s key-value cache, increasing throughput
[1]https://modal.com/docs/examples/text_generation_inference
Therefore, unsurprisingly, the cost per item of inference on a batch of items is significantly lower when the batch is e.g. 8 than 1 (in the case of Transformers there are further gains to be made because roughly half of the attention calculations in token k+1 are identical to the calculations of token k and can be easily reused by writing the formulas a certain way, the keyword to look for is causal attention mask).
In any reasonable GPU inference setup the weights would be preloaded.
The main reason I can come up with why doing the same calculation 8 times in parallel instead of 8 times sequentially is that you benefit from better locality of reference.
> ANPD dynamically generates draft outputs via an adaptive N-gram module using real-time statistics, after which the drafts are verified by the LLM. This characteristic is exactly the difference between ANPD and the previous speculative decoding methods.
ANPD does provide a more general-purpose solution to drafting that does not require training, loading, and running draft LLMs.
I don't think it would be that hard to switch it out for a pretrained ngram model.
Especially when it comes to programming languages or data formats like json - a lot of compute is spent on ie. silly whitespaces.