Honey, I shrunk the embeddings: Matryoshka vs. PCA
dylancastillo.co
dylancastillo.co
> You can push this further by combining quantization with truncation or PCA. The resulting vectors can be dramatically smaller while still preserving a surprising amount of retrieval quality.
Counterintuitively - quantisation can also be combined with a random rotation step before the quantisation. A random rotation spreads information across more dimensions, allowing more aggressive quantisation without losing accuracy. Ironically - almost the opposite of a PCA.
I do wonder if relevant here though. It relies on the embeddings having "structure", i.e. that principal components point along basis vectors, which may not be the case with text embeddings.
Source: https://research.google/blog/turboquant-redefining-ai-effici...
sometimes adding noise over a frequency range is better than removing it entirely, the opposite tends to be true for text and especially code, where you'll want lists of language and framework keywords, and then a pass on top of that to scan the codebase itself for its own slang. You can then 'double dip' and use these lists to weight the results after the fact
In my experiments, I used lots of embedding models and the results were not nearly as uniform as this curve, just FYI. I didn’t use any of the API-based models though
I also wrote about this exact comparison when using PCA and MRL to quantize static models, see: https://stephantul.github.io/blog/mrl-pca/
I couldn't find much when I first looked into this, which is why I ended up writing the article.
So I guess the answer is: no
Feels like a “just use logistic regression” moment :)
That is where something like Matryoshka embeddings has appeal. You trade a little bit of performance for a guarantee of training + validation set coverage.
In your conclusion, you report that PCA won on most dimensions. Did you investigate why you found that PCA outperforms MRL when the original paper found that their SVD baseline did not?
Are the goals different, or should the original paper have done something more similar to your benchmark? Or something else?
This is no way a new problem. The idea that embeddings or even search didn't exist until LLMs is simply false.