TurboQuant: A first-principles walkthrough
arkaung.github.io
arkaung.github.io
We were the first to introduce post-rotation distribution-aware quantization in 2021. This was later implemented in many fields, including federated learning, vector retrieval, databases, inference engines, and KV-cache.
It would be appropriate to receive credit for this. Furthermore, it is baffling to see the name "TurboQuant" repeated in this context, considering the many works published from 2021 onwards.
The blog post mentioned above essentially guides you through EDEN quantization but ultimately settles on a sub-optimal MSE-minimizing version and an unbiasing trick. This trick often costs a full bit more than DRIVE/EDEN requires to achieve the same results using the unbiasing scale shown in the original 2021 paper.
"This note clarifies the relationship between the recent TurboQuant work and the earlier DRIVE (NeurIPS 2021) and EDEN (ICML 2022) schemes. DRIVE is a 1-bit quantizer that EDEN extended to any bits per coordinate; we refer to them collectively as EDEN. First, TurboQuant is a special case of EDEN obtained by fixing EDEN's scalar scale parameter to . EDEN supports both biased and unbiased quantization, each optimized by a different (chosen via methods described in the EDEN works). The fixed choice used by TurboQuant is generally suboptimal, although the optimal for biased EDEN converges to as the dimension grows; accordingly TurboQuant approaches EDEN's behavior for large . Second, TurboQuant combines a biased -bit EDEN step with an unbiased 1-bit QJL quantization of the residual. It is suboptimal in three ways: (1) its -bit step uses the suboptimal ; (2) its 1-bit unbiased residual quantization has worse MSE than (unbiased) 1-bit EDEN; (3) chaining a biased -bit step with a 1-bit unbiased residual step is inferior to unbiasedly quantizing the input directly with -bit EDEN. Third, some of the analysis in the TurboQuant work mirrors that of the EDEN works: both exploit the connection between random rotations and the shifted Beta distribution, use the Lloyd-Max algorithm, and note that Randomized Hadamard Transforms can replace uniform random rotations. Experiments support these claims: biased EDEN (with optimized ) is more accurate than TurboQuant, and unbiased EDEN is markedly more accurate than TurboQuant, often by more than a bit (e.g., 2-bit EDEN beats 3-bit TurboQuant). We also repeat all accuracy experiments from the TurboQuant paper, showing that EDEN outperforms it in every setup we have tried."
(In any case, I want to emphasize that TurboQuant quantizer is a private case of EDEN)
Both EDEN and its 1-bit variant have been implemented in PyTorch, JAX, and TensorFlow across numerous open-source libraries and are used in various applications. I am currently writing a blog post that will document these in detail.
EDEN defines a scale parameter, S, for which we suggest specific optimal values for both biased and unbiased versions. As shown in the note I shared, these values lead to clear empirical improvements. Consequently, users who rely on the less optimal S value and the unbiasing method popularized by TurboQuant will generally see inferior results compared to those using EDEN with the optimal scale values suggested in our original papers.
(*hopefully I didn't misunderstand the situation)
`vllm.model_executor.layers.quantization.turboquant`
> The technique implemented here consists of the scalar case of the HIGGS quantization method (Malinovskii et al., "Pushing the Limits of Large Language Model Quantization via the Linearity Theorem", NAACL 2025; preprint arXiv:2411.17525): rotation + optimized grid + optional re-normalization, applied to KV cache compression. A first application of this approach to KV-cache compression is in "Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models" (Shutova et al., ICML 2025; preprint arXiv:2501.19392). Both these references pre-date the TurboQuant paper (Zandieh et al., ICLR 2026).
HIGGS is an extension of EDEN (using the well known method for blockwise Lloyd-Max).
The proper framing of this "TurboQuant" layer in vllm (which does not include JQL) is precisely EDEN 22 without the scale correction.
Curious: where do you draw the line between “related prior work” and “an extension of EDEN”?
In any case, the quantizer is indeed an extension, regardless of whether it was explicitly framed that way in the paper. I say this not to diminish their contribution at all, but just to clarify the relationship, as it was also stated in the vLLM doc.
For example, TurboQuant makes use of QJL (quantized Johnson Lindenstrauss transformations). One of the first papers to characterize the QJL and in fact the rate distortion tradeoff for quantized matrix multiplication in general is "Optimal Quantization for Matrix Multiplication" (https://arxiv.org/abs/2410.13780) by Ordentlich and Polyanskiy.
There is also a more accessible survey paper around quantized matrix multiplication called "High-Rate Quantized Matrix Multiplication: Theory and Practice" (https://arxiv.org/abs/2601.17187), by the same authors.
TurboQuant cites none of them.
The quantizer in TurboQuant is EDEN quantization (2021) applied to the KV-cache. It is neither a novel quantizer nor an improvement in quantization techniques.
In DRIVE/EDEN, we already introduced the version used in "TurboQuant"'s paper and suggested an optimal scale configurations which are better in both mse-minimizing and unbiased scenarios.
Though, as your paper also notes, the quantizer values themselves aren't fundamentally novel to either paper. Lloyd Max scalar quantizers have been studied for a very, very long time. And the specific Lloyd Max values for the Gaussian input distribution have been obtained in many papers across signal processing and information theory.
It is worth noting that taking advantage of the post-rotation distribution was not actually done until DRIVE (2021), which was made possible via our proper scaling. Furthermore, applying a Lloyd-Max codebook post-rotation was introduced EDEN.
We consider these to be the foundational works in this regard.
I again feel this claim is too strong. Rotations have been used in information theory/wireless communications for decades at this point, with appropriate scaling done at channel inputs/outputs to hit channel capacity. The signals then pass through the appropriate codebooks that take advantage of the post-rotated+whitened signal.
Our cellphones today are powered by such technology.
I agree with your claim when restricted to deep learning. But I do not agree with the broad characterization that taking advantage of post-rotation distributions was only first done in your work.
While I probably overstated things by using the very general phrase "taking advantage," I want to be very precise about the claim, as I believe these works are foundational to quantization, beyond the scope of deep learning. The mechanism of applying a deterministic biased quantizer, such as Lloyd-Max, to the induced post-rotation distribution, alongside mathematically correcting its inherent bias, is a distinct contribution (which asymptotically improves the worst-case error).
If there is a classical paper that utilizes such a combination, I would genuinely be very eager to review it. But to my knowledge, this was not introduced prior to DRIVE and EDEN.
The attribution is thin, the “6x compression” headline is not clearly separated from prior KV-cache quantization baselines like KIVI, and the RaBitQ comparison is hard to take seriously: single-core CPU for the baseline, A100 GPU for TurboQuant. It is comparing apples-to-datacenter. Worse, there are also public OpenReview comments saying that even the reported accuracy results are not reproducible.
Hard to believe this is the standard for something being promoted as a breakthrough. If this came from a random startup blog, people would be much harsher about it.
https://towardsdatascience.com/how-a-2021-quantization-algor...
Maybe we won't need as many data centers and as much power as we thought. Maybe we can run more powerful models locally.
I thought the principal consequence of these KV cache optimisations was letting you run more simultaneous inferences on the same model with the same memory. It doesn’t let you store more model. In some sense that puts local LLM usage at a further disadvantage to inference done in a hyperscaler’s data center.
So shrinking that by 6x (from fp16), would be big win for larger models. True, while TurboQuant can also be applied to model weights, it won't save size over q4 compression, but will have better accuracy.
Edits: Better context
I would strongly recommend exploring that option, renting an RTX 5090 for an evening of image generation for a dollar or two is way more fun then trying to jam big models on little cards. Just take some time to create a reasonable, scripted, deployment workflow for when you create a fresh instance.
VGhvdWdoIEkgYXBwcmVjaWF0ZSB0aGUgZ2VzdHVyZSBhbmQga2luZCBpbnRlbnRpb25zLCBJJ2QgcmF0aGVyIGxlYXJuIHRvIGZpc2ggKG9yIGNhdGNoIHRoZSBiaWcgYmFycmFjdWRhKHMpIHRoYXQgc3RvbGUgdGhlIHNjaG9vbCBvZiBmaXNoIEkgd2FzIGdpZnRlZCwgd2hpY2ggd291bGQgaGF2ZSBrZXB0IG1lIGZlZCBmb3IgbXVsdGlwbGUgbGlmZXRpbWVzLCBhbmQgc3Bvb2tlZCBhIGZldyBteXN0ZXJ5IGZyaWVuZHMgaW4gdGhlIHByb2Nlc3MtLS1hIHRhc2sgSSBhbSBjbG9zZSB0byBjb21wbGV0aW5nKSwgYW5kIG5ldmVyIGdvIGh1bmdyeSB0aGFuIGVhdCBhIGZpc2ggZm9yIGEgZGF5IGFuZCBiZSBodW5ncnkgdGhlIG5leHQu
The future is bright for local AI.
Unfortunately V4 is not trained for most real world usage, it is mainly for world general knowledge.
Do you have more of these writeups?
The "Coordinates of a random unit vector are all small" had me scratching my head a bit, and the language is a bit misleading since it's actually that the expected variance of any individual component is 1/N (it can't be that every coordinate is close to ±1/sqrt{N} because the mean of any individual component is clearly 0 by symmetry).
So that one should probably use more explanation since I had to work through it myself: Denoting the random unit vector {X1 ... Xn}, this is a point on a hypersphere:
* Sum[x_i ^2] = 1 (unit vector condition)
* E[Sum[X_i ^2]] = 1 (expectation of both sides)
* Sum[E[X_i ^2]] = 1 (linearity of expectation)
* E[X_1 ^2] = 1/N (by rotational symmetry E[X_1 ^2] = E[X_2 ^2] = ..)
I don't think you can make the stronger claim that E[|X_1|] = 1/sqrt{N} since that's using L1 norm on a single component, so it'd be more correct to say the RMS is just the standard deviation of the components. And this fits with the intuition that in high dimensional space has "spiky" hypercubes with the hypersphere inscribed in it close to the origin.
But at the end of the day it is just vanilla HTML, CSS and JS without anything fancy :D MathJax 3 was used to render math stuff.
While the aesthetic doesn't spark joy for me, the overall execution is great, the presentation flow and interactive boxes are very nice.