SimSIMD: Hardware-accelerated similarity metrics and distance functions
github.com
github.com
Does this do levenshtein distance? Didn't see it in the main document.
Edit: also, THANK YOU for adding a parameter to specify # of threads. Wish this was a wider standard practice.
Oh, and for multithreading, USearch might be more practical.
If the first 4 bytes of the string are the same, the strings are likely to be equal. Similarly, the first 4 bytes of the strings can be used to determine their relative order most of the time.
That's a fun heuristic.L2 in low dims, Cosine in high dims. Hamming on similar size sets, Jaccard on others. Jensen-Shannon (symmetric KL) when dealing with histograms and probability distribution.
So really the answer honestly tends to be ad hoc: "whatever works best". It's good to keep in mind that any intuition you have about geometry goes out the window as dimensions increase. It's always important to remember assumptions made, especially when focusing on empiricism. There are definitely some nuances that can point you in better directions (pun intended :) than random guessing, especially if you know a lot about your geometry, but it is messy and nuances can make big differences.
I wish I had a better answer but I hope this is informative. Maybe some mathematician will show up and add more. I'm sure there's someone on HN that loves to talk about higher dimensional geometry and I'd love to hear those "rants."
And for me to get a feel for it.
The "works best" is also, in many cases, subjective. This is also not easy to assess, you may need several people looking at the results, here molecules, to say if yes they are similar or not. A chemist will think differently than a biologist in this regard.
[0]: https://www.chemeo.com/similar?smiles=Cc1c%28%5bN%2b%5d%28%3...
[1]: https://jcheminf.biomedcentral.com/articles/10.1186/s13321-0...
C11 is also a bit tricky, as its support is optional, as far as I remember.
Also interesting that it beats NumPy/SciPy by so much! I wonder what they're doing..
https://github.com/scipy/scipy/blob/main/scipy/spatial/src/d...
https://github.com/facebookresearch/faiss/wiki/How-to-make-F...
There are more materials in my blog, including 4 articles from October 2023: https://ashvardanian.com/archives/
Benchmarks are quite easy to reproduce, especially in Python. The numbers for AWS Graviton 3, Apple M2 and Intel Sapphire Rapids CPUs can be found here: https://ashvardanian.com/posts/simsimd-faster-scipy/
Multiply it by a more efficient HNSW implementation and you might be looking at a 100x performance difference on a scale of a billion vectors: https://www.unum.cloud/blog/2023-11-07-scaling-vector-search...