Fantastic results. Well done.
...So this is built into the way the model works.. if I'm understanding it correctly.
I was wondering what would be involved in getting it to work with GGUF files, rather than safetensor files...
I was wondering what would be involved in getting it to work with GGUF files, rather than safetensor files...
At the moment not even MTP is merged into llama.cpp, so I wouldn't quite hold my breath for it.
Hope the paper gets lots of references and the technique gets a lot of use to save power and time.
There's been several potential big changes for LLM inference efficiency over the last few months. There's been Attention Sequencing (I think it's called..?) Turbo Quant and this one.
Interesting times.
In the meantime I've benchmarked Orthrus some more and got some quite promising results. So I'd be glad if my prediction that it may take some time until it lands in llama.cpp turns out to be wrong.