But I don't see them when I filter the list for 'voyage'.
But I don't see them when I filter the list for 'voyage'.
Many of the top-performing models that you see on the MTEB retrieval for English and Chinese tend to overfit to the benchmark nowadays. voyage-3 and voyage-3-lite are also pretty small in size compared to a lot of the 7B models that take the top spots, and we don't want to hurt performance on other real-world tasks just to do well on MTEB.
Why should I pick voyage-3 if for all I know it sucks when it comes to retrieval accuracy (my personally most important metric)?
Nice!
Fortunately MTEB lets you sort by model parameter size because using 7B parameter LLMs for embeddings is just... Yuck.
It is worth noting that their own published material [0] does not entail any score from any dataset from the mteb benchmark.
This may sound nit picky, but considering transformers' parroting capabilities, having seen test data during training should be expected to completely invalidate those scores.
[0] see excel spreadsheet linked here https://blog.voyageai.com/2024/09/18/voyage-3/
Could hurt performance in niche applications, in my estimation.
Looking forward to try the announced large models though.