eshoyuan··on Lossless Acceleration of LLM via Adaptive N-Gram Parallel DecodingI don't think this will be better than weaker transformer.
eshoyuan··on Lossless Acceleration of LLM via Adaptive N-Gram Parallel Decodinghttps://github.com/ggerganov/llama.cpp/tree/master/examples/...