Challenges in Implementing a Full-Text-Search Engine
bhavaniravi.com
bhavaniravi.com
FTS is still commonly included in e-commerce where I guess it's not quite dead but only because truly relevant search results are irrelevant to retailers' bottom line.
SS is at least one magnitude more complex to set up, especially if you want to be able to refresh your "index" and even more so if your data is big. What a typical back-ender would wip up in a couple of days using ES/Solr you now need a proper Math Dev to do and for the MD's model to become useful you need them to hand it over to a distributed systems expert.
SS through ML is commonly a nasty, duct-taped work-flow that at best results in a system that looks more like a POC or DEMO than a proper production system, unless you work at Bing/Google that is (probably, but I have no hands-on experience of those systems).
I've been trying for years, I'd say at least ten, to try to "commoditize" this work flow, make it simpler, more usable for generalist devs (not just ML people) but no matter what I do I keep getting crushed under the weight of the data. To me it seems search is dead and we haven't appointed a new king.
Very high performance, and much deeper support for tensors than elasticsearch has with with its new dense_vector type. Sadly they both have the same "gotcha" in that it is only really practical to run any _decent_ ranking model as a second phase re-rank to avoid linear scans.
I agree semantic search is useful, but what you're proposing sounds vague, like black box magic.
To provide semantic search, you still do the dirty things you just mentioned: integrate a stemming library, integrate synonyms and a huge corpus of 2+ words topics (ie. when someone searches for "big data", you should always return documents with those terms together, and never apart). You need those things. ML might help you generate them, but it isn't magically going to tell you X, Y and Z documents should also be returned for a query, even though it doesn't contain that term.
This in itself is a good enough reason to be a sufficient argument to refute your main claim that FTS is dead. It will not be dead until an alternative is sufficiently commoditized so that regular backenders can set it up. It may be edged out of certain use cases where it is mission critical to get fantastic results, but there are MANY places of use where good old boring FTS is plenty good enough. Otherwise it wouldn't have been used in the first place.
While I agree that building a great and relevant search experience with a Lucene-based engine requires lots of extra time and effort to get right, there are other non-TFIDF based solutions that provide a much faster path to great relevance with far less effort (https://blog.algolia.com/inside-the-algolia-engine-part-1-in...), and it's possible to have semantic ranking without too much machine learning (https://blog.algolia.com/promote-search-results-query-rules/). Not to discount the value of machine learning - we're finding that for specific usecases ML can be a very valuable way to help surface more pertinent content for individuals based on their profile/preferences etc. (https://blog.algolia.com/personalization-announcement/).
This may be along the lines of what you mentioned around "commoditizing" complex traditional search workflows. I'd be curious to hear more about what kind of use-cases you think are trickiest without SS.
I used to be quite impressed with Lucene, even at version 1.0 (when a fuzzy search meant a full table scan), then watched in joy when they conquered the search market, before realizing how it struggled (and still does) with, well, I hate to say it (because I'm usually ridiculed when I bring this up), y'know, big (-ish) data. The proposed and popular solution: sharding the data onto a cluster of machines.
Algolia seems to be a focused, streamlined and more efficient ElasticSearch, at least in the FTS use case.
I've worked almost exclusively in e-com for ~20 years. Algolia FTS+personalization seems to fit the e-commerce use case pretty darn well.
I wonder, regarding "Algolia Query Rules" (which also seems like a real killer-feature for e-commerce):
>> automatically transforming a query word into its equivalent filter (“cheap” would become a filter price< 400...
How do you translate "cheap" into "price<400"? By maintaining a dictionary? Also, what if some people think 400 is quite expensive?
I want to build or implement a search engine that is inherently self-maintained in the same way you and me are self-maintained. As humans, however, we do have a serious flaw. In order for us to maintain an index of our knowledge we need to sleep. To start with I'd like to try to mimic that construct, then move past it.
This is being worked on in Solr community. See an in-progress book: https://www.manning.com/books/ai-powered-search And related issue for an example of work: https://issues.apache.org/jira/browse/SOLR-9418
There has been some work but not sure when it will be stable, needs a new kind of index:
https://www.slideshare.net/lucenerevolution/what-is-inalucen...
These two features would allow you to do visual search, semantic text search, recommendations, learning to rank and etc.
We've frequently had the same dream of adding more native support for nearest-neighbor type queries, since that is the workhorse of so many useful techniques in the modern NLP stack.
Right now, we have lots of dense vectors stored in massive toast tables in PG. It's faster to fetch them rather than recompute them, especially since there are a number of preprocessing steps that limit what we pay attention to.
The discussion here about full text search versus semantic search is interesting. In our experience, both are highly relevant. Sometimes it's most useful for our customers to segment their conversation data by exact text matches, and other times semantic clustering is most effective. I think there's plenty of reason to offer both kinds of capabilities.
https://www.elastic.co/guide/en/elasticsearch/reference/curr...
The vectors are only used for scoring, not matching, but they are working on a ANN model for that.
https://github.com/avremel/lucene
Elastic/Solr is a very decent option. Last time I checked, Algolia and other SaaS were too expensive for small businesses.