If this is "plug and play" can it be added to say llama.cpp and give ~3.67 speedup to existing models or is there some complication?
> ANPD dynamically generates draft outputs via an adaptive N-gram module using real-time statistics, after which the drafts are verified by the LLM. This characteristic is exactly the difference between ANPD and the previous speculative decoding methods.
ANPD does provide a more general-purpose solution to drafting that does not require training, loading, and running draft LLMs.
I don't think it would be that hard to switch it out for a pretrained ngram model.