Building a Content-Based Search Engine II: Extracting Feature Vectors
deepideas.net
deepideas.net
I've build an interesting NN-based model, and I've been thinking about using some of the early layers as feature for a search engine.
The obvious thing to do is to dump the feature vectors into Elastic or Solr or something.
Ideally what I want is to:
1) Put 1024 dimensional vectors of floats into the index
2) Use plain cosine distance as the distance metric
2b) Customize the distance metric (or preferably use a preexisting and optimized earth-mover-distance implementation).
My initial Googling indicated that 1 and 2 are harder than I expected - it seems both Elastic and Solr don't have good representations for vectors, and assume you want BM25 or TF-IDF for your ranking.
Surely I'm missing something? Not super-keen on having to drop back to using Lucene.
Also https://erikbern.com/2018/02/15/new-benchmarks-for-approxima...
You can use something like annoy for searching vector space. It's not made specifically for text search, so you'll have to roll your own normalization and so on, but that's usually the less complicated part anyway.
I read https://erikbern.com/2018/02/15/new-benchmarks-for-approxima... (by the author of Annoy) with great interest.
He points to HNSW which I've know nothing about, but seems very fast!
What does querying it look like?