What is full-text search (and why you need it now)
blog.indextank.com
blog.indextank.com
It's common that I get search results that don't contain the keywords I searched for. The results will also contain the keywords but in different parts of the page, so short phrases like hacker news will match pages with "hacker" in the first paragraph and "news" somewhere else on the page. Google used to recognize a hyphen as meaning "only return results where these two words are joined by a space or puncuation", now it ignores the hyphen and you have to enclose the phrase in quotes to get that behavior.
Even more interesting than a simple full-text search would be a regular-expression full-text search of the internet.
In either case, I'm happy to hear IndexTank is trying to fix this.
http://www.googleguide.com/wildcard_operator.html
UPDATE: and someone just posted about the AROUND() operator, http://news.ycombinator.com/item?id=1983930
It would seem to me that indexing data in such a way as to support arbitrary regex would be difficult. Anybody know of an existing system?
If you type "cat AROUND(5) dog" into google, it'll return the top results that contain anywhere from 0 to 5 words between cat and dog.
I think a fast full-text search for the desktop can be very useful. Collecting all those publications doesn't really help when you can't find them when you need them.
But most importantly, you don't have to deal with installing, configuring, managing and scaling a sphinx setup. IndexTank is cloud-based, you don't have to deal with servers, RT indexes, hadoop, and so on. We worry about all that.
I had a thought that relevance should be computed like a "stream database", but with an index rather than a data store behind it. Maybe something like Streambase - but written in Clojure - on top of the lucene index.
Take it one step further, where the entire index is expressed as s-expressions (i'm a bit out of depth here) - you can basically write a mapreduce job on the index.
However, I'm not so sure that computing relevance can be parallelized - i.e can it be broken down as opposed to be computed as a global problem ?
http://indextank.com/documentation/function-definition
(Feedback on this documentation is more than welcome, it's a work in progress!)