Full Text Searching with Solr and Sunspot
collectiveidea.com
collectiveidea.com
http://freelancing-god.github.com/ts/en/
CollectiveIdea also maintains a good fork of delayed_job for those who are interested in that.
EDIT:
- sunspot allows for per-request weighting in the search (see http://blog.logeek.fr/2011/3/4/hackerbooks-books-stackoverfl... for an example), not sure if it's possible with ts
Yeah, if I ever use a NoSQL datastore on a project I'll likely end up using Solr. Faster than writing a new Sphinx interface for it.
That, and I had a terrible experience using acts_as_solr back in the day...
I do have to disagree with you on ThinkingSphinx. I have used both, and I gel much better with Solr. Don't get me wrong, the guy(s) who did ThinkingSphinx have worked hard. I just don't like the moving pieces of ThinkingSphinx: continuous rake tasks, the fact that it sits directly on top of the database, extra re-deployment steps when using capistrano. I prefer having the search server sitting apart.
Different strokes for different folks. :)
a) developing the 2 or 3 indexes you need for good search results b) running test suites c) production
That said, in the past, I haven't been able to transfer the index configuration from sphinx to SOLR indexes and get 98% matches between the 2 engines. Sphinx (thinking sphinx in rails) used to do funny stuff in the pre-1.0, like if you put in 2 search terms you would get X results, if you put in those 2 terms plus another, you would get more results. I think they've fixed most of that in 1.0
Create a counter table to keep track of what records you've indexed and cron frequent indexer runs using the latest records as offsets.
If you have documents that frequently change, maintain your full index less often and instead merge the two on a more regular basis.
Sphinx's indexer speed is one of its advantages, but it's largely dependent on the efficiency of your SQL and underlying MySQL indices. Perhaps something else is indirectly influencing your indexing performance.
The indexing speed is so very fast-- I can't say enough good things about sphinx.
1.'did-you-mean' support
2. facets
3. replication
4. spell checking
5. auto-complete is easier to implement with Solr
6. Sphinx doesn't return whole documents, only the doc id
7. Sphinx doc ids must be integers