Full text search indexing stemmer added to MongoDB core
jira.mongodb.org
jira.mongodb.org
I don't know what the plan is, but an integrated stemmer alone isn't of big value to me. Actually I prefer to have a stemmer in my own code, so I can tweak/update it without touching a database.
that's a bad news for users of all those PHP/RoR/Django/node etc. apps, who will never get proper on site search functionality. majority of lazy devs won't go for Solr-like solution
What exactly is the issue/problem that you're trying to convey?
One can use any search solution, from the most basic to the most advanced one, with any server-side technology.
Even if your hosting provider doesn't allow you to run some technology stack, there are hosted search solutions with APIs you can use.
And it's not like any but the most basic of sites should not use at least a VPS anyway.
http://stackoverflow.com/questions/9160305/elastic-search-vs...
http://adventuresincoding.com/2012/05/full-text-search-in-ra...
http://www.slideshare.net/dkeener/rails-and-the-apache-solr-...
I don't see the bad news at all, if you wan't to implement proper search you need something like that.
anyway, what if my hosting provider won't let me run Solr or any other java software?
Nothing in PHP/RoR/Django/node.js makes them incompatible with Lucene and/or Solr. You just need to run a jvm in parallel.
And it's not like every page needs a "full text indexing and search" solution.
Personally, for a lot of use cases I prefer exact string matches over BS stem indexing.
Really? I've worked on a few search projects in different spaces (venues (aka places/stores), source code, and products) in the past, and while exact string matches are often a good sign of quality, stemming and other analyzers make huge improvements in recall (and when measuring transaction volume in A/B testing strict string matching performed substantially worse). Certainly if you throw out the exact match signal (i.e. only index stemmed) I've seen that result in a deterioration of quality. What sort of data do you work with?
The intention was to quickly develop the extra search pieces needed in Python and then port them to solr. (For example we needed a custom scoring mechanism, and needed to experiment with spelling errors, pronunciation equivalency etc). However Whoosh turned out performant enough that I didn't need to touch solr again (XML config files always make me judder!)
So if you Python, I strong recommend giving Whoosh a go especially when starting out a project as you'll be more productive.
from an impl pov I suspect the stemmer is the (logical) next step (of many) on this long road of implementing full text search features.