Image-Match: Open-source scalable reverse image search
github.com
github.com
FWIW, John Resig uses pastec for his work:
http://ejohn.org/blog/image-similarity-search-wanted/
http://ryanfb.github.io/etc/2015/11/03/finding_near-matches_...
Here's a real-world use case http://blog.livefyre.com/architecting-sidenotes/
We used MoreLikeThis to reduce our queries count to the 30-40 most statistically interesting terms. The one hiccup being an issue in Lucene [1] where the term cache wasn't operating properly. We just added our own image query term cache and a custom MLT query to leverage it, which gave us a 10x speed bump over any other methods we tried.
The interestingness of the terms is assessed on a per-term basis though, so you might see a relevence drop for some types of image if you set MoreLikeThis to use too few terms.
Fortunately or unfortunately, we were already achieving pretty good speed with Elasticsearch so we didn't implement it. However, it didn't occur to me to try a MoreLikeThis query, which should be even simpler -- I will look into it!
I tried something similar; but with a different approach. I tried creating compound words, a bit like n-grams. I didn't get it working as that was a side-project and I couldn't commit enough time.
Does anyone know if this has changed since?
Rest assured, this is the correct one!