Caponia - an in memory, full text search system in clojure
github.com
github.com
However, (at least via couchdb) Lucene needs a whole JVM of its own, and I'm planning to run this on a VPS with pretty tight resources. A couple of MB in the same JVM as my site is a whole lot more appealing, and has the additional advantage of using native data-structures directly.
Maybe you were thinking of Solr? That's the REST service that wraps Lucene but you can use the Lucene API from within your app just as easily if that fits your use case.
On the other hand, I'm quite happy with how this turned out, it's one less dependency, and it's been fun to write.
If you are using Clojure, Scala, JRuby, etc., might as well take advantage of great Java libraries.
Is that all that's needed to implement the full text search?
And it's fast?
The particular pathological case for this system is query where the search stems occur in the majority of documents (there is a small exclusion list that helps, but its not a panacea) because the merging of stem results is at best a linear operation (:or queries are currently much simpler than :and queries)
The other difficult case for this system is updating or removing a document from the index; removal is fairly easy: generate the stems for the current state and update those stems in the index.
To update (an existing document in the index) efficiently you need the set of stems in both old and new versions of the document; For the stems that are in new set you can just add them to the index as normal. For the difference of the old set and the new set (ie, those stems only in the old set - clojure's difference function is not symmetric i believe) you then need to remove the document id from those stems in the index.
edit: clarified meaning of update.