KVarN: Native vLLM backend for KV-cache quantization by Huawei
github.com
github.com
Am I reading this right??
But the point was that quality didn't magically increase.
edit: It might not be clear that it is based on vLLM 0.22, which is the current version: https://github.com/huawei-csl/KVarN/commit/d6290e99098d7426d.... All you have to do is create a diff off it; it's fairly straightforward.