It’s not meant to be a state of the art model though. I’ve put in pretty limiting constraints in order to keep dependencies, size and hardware requirements low, and speed high.
Even for a word embedding model it’s quite lightweight, as those have much larger vocabularies are are typically a few gigabytes.
I think this one is currently the top of the MTEB leaderboard, but large dimension vectors and a multi billion parameter model: https://huggingface.co/nvidia/NV-Embed-v1
A lot of people wind up using models based purely on one or two benchmarks and wind up viewing embedding based projects as a failure.
If you do answer some of those I’d be happy to give my anecdotal feedback :)
As of the last time I did it in 2022, Mini-lm can be distilled to 40mb with only limited loss in accuracy, so can paraphrase-MiniLM-L3-v1 (down to 21mb), by reducing the dimensions by half or more and projecting a custom matrix optimization(optionally, including domain specific or more recent training pairs). I imagine today you could get it down to 32mb (= project to ~156 dim) without accuracy loss.