I don't need exact results. FP8 quantization is almost lossless and even 6-bit quantization is usually acceptable. Can this be combined with quantization?
It ends up being somewhat faster than regular speculative decoding in normal setting (GPU only). If you are doing CPU offloading it's massively faster.
Edit typo
It is in their TODO part in https://github.com/Infini-AI-Lab/Sequoia/tree/main