A fast, fuzzy, full-text index using Redis
playnice.ly
playnice.ly
What's better? Use a tool that's optimized for the job. I would recommend looking at Sphinx, which is quite amazing and can handle indexing billions of words on a single server without using much CPU or memory. Plus Sphinx has tons of features, such as geo-based search, fuzzy search, boolean search, full support for Unicode etc. etc.
You could probably extend this to include that kind of thing with some clever hacking, but as it stands this is still a long way from competing with something like Sphinx. Still - it's a nice little demo of using Redis for something and taking advantage of the specific strengths of Redis such as first class set support.
A larger, Redis-backed search engine would certainly be an interesting project though...
redis.zincrby(metaphone_key, item.item_id, 1)
This will allow you to sort items by highest number of occurrences
I certainly agree that a purpose-built solution could be better if you needed the features you mention, but that was not our use-case. This provided an elegant solution which would avoid us having to duplicate all our item metadata. Also, given our schedule it seemed clear that Redis VM would be available by the time we needed it.
It is true that Redis is not optimised specifically for full-text search. However, it is optimised for serving data structures very quickly which is exactly what we needed.
I will probably go with sphinx for it's other capabilities but this sounds better than the other redis search implementations that are out there.
Btw. like I mentioned Sphinx does support fuzzy search (and boolean search), so I don't really know what you win by rolling your own solution in Redis, other than worse scalability and a crippled full-text search.
Could we have learnt and used another technology? Sure, and we may still do so. For now this solution works well for us.
Should other people use this technique? Like always, it depends on the situation.
Solr has effectively the same feature set, including geospatial and multilingual searching, but it has better relevance, and generally is faster at returning query results (although Sphinx is faster at doing full reindexes).
Also, unlike Sphinx, Solr doesn't glue you to MySQL. Or to any SQL database; use it with Cassandra, Voldemort, text files, quantum storage in the galactic hive-mind, whatever.
(edit: They've added PostgreSQL support, and raw XML support, since I last installed and tested Sphinx about a year ago)
Plus, unlike Sphinx, you don't need to reindex when you add new records, because Solr can seamlessly merge indexes on-the-fly.
I know I sound like a fanboy here, but I spent a lot of time evaluating the two of them for our product, and Sphinx just didn't fit the bill for a large number of reasons. It's a good solution if you just need fulltext indexing in a MySQL database, but if you want to move beyond that, have a look at Solr.
http://beerpla.net/2009/09/03/comparison-between-solr-and-sp...
We're replacing our text indexing solution with Solr and it's been a phenomenal transition.
EDIT: it's fine if your items are very small or do not change often
Also, for our use case it is not necessarily a bad thing that removed words may cause an item to be shown in the results. After all, just because a paragraph was removed from a bug description doesn't necessarily mean it is invalid search fodder (bearing in mind that, in our case, full-text search is just one way of filtering).
Updating the index should be a simple as calculating the old metaphones, removing from the index, and then reinserting the key based on the metaphones of the new content (multi/exec would is perfect for this)
And from what I have read of Sphinx it likely has the same issues. Plus it seems to be aimed only at the lamp crowd - not very friendly if you are not using php+mysql already.