Neural Search for medium sized corpora
github.com
github.com
https://news.ycombinator.com/item?id=29551947 - About semantic search (ie, vector search applied to text)
https://news.ycombinator.com/item?id=29554986 - About vector search at Google
https://news.ycombinator.com/item?id=29555780 - Open-source vector search index from Facebook
https://hanxiao.io/2018/01/10/Build-Cross-Lingual-End-to-End...
The other stuff on that guy's (Han Xiao's) blog also looks interesting. I think this neural search stuff explains why we're all getting such low-precision results on Google/DDG by the standards we were used to. But, the claimed advantages are impressive. This is a big change in how search engines work, and I was unfamiliar with it before.
I still have reservations about where enough training data is supposed to come from, without already having a big system and user base. But, maybe some more info about that is out there. So this is another thing to look into. Wow.
Search implements a wrapper of the Python ElasticSearch client that is scalable and dedicated to corpora composed of tens of millions of documents.