Its like the SQLite of search. Single library that provides the basic set of features you would expect in a quality search experience: facets, ranked search, boolean operators, stemming etc etc.
Its like the SQLite of search. Single library that provides the basic set of features you would expect in a quality search experience: facets, ranked search, boolean operators, stemming etc etc.
If you’re using the JVM already and want the SQLLite of search, you can instantiate an index in your own JVM, load documents, and get compatible query features.
You can also use Lucene directly, which is simpler, but does have some nuances that differ if you want perfect query compatibility between a deployed ElasticSeach cluster and an in-memory search.
Elastic the company “supports embedded” the same way it “supports self hosted”. It’s left as an exercise to the user.
https://www.djcbsoftware.nl/code/mu/
https://manpages.debian.org/buster/maildir-utils/mu-query.7....
I was previously looking into embedding TypeSense somehow (but not portable), MeiliSearch (but their API triggered a panic response), or SQLite's full-text search extension (but limited features e.g. language support). Using a dedicated library would be much better and this seems to be packed with features.
Could you expand on that with examples? I only played around with it lightly but it seemed ok via the Python SDK.
I've been recently leaning towards Meilisearch for some client facing search features. This was over Typesense because, i believe Typesense is totally in-memory, but in some cases my dataset might not fit in RAM(atleast not at a reasonable cost). I was fine with taking a performance hit for disk access, provided i could have thousands of separate indexes with separate permissions etc. and directly call them from the browser app.
* If you have more than 65,535 words in a document, words after that limit will be SILENTLY IGNORED (not indexed). This was 1,000 words when I tried it, which was very limiting
* Documents where your query word happens sooner are ranked above documents where your query words happen later (in the document string). You can't turn that off
* You can't get the "match" information (i.e. the position where your query matched in the document) unless you retrieve that attribute (i.e. tell the database to send you the entire document content)
* The API uses HTTP verbs wrong, e.g. POST is used to "Add or replace documents" and PUT is used to update documents (it's exactly backwards from the intended use, per RFC 7231)
This discussed more at https://trac.xapian.org/wiki/Licensing
To be clear, I don't hate the GPL license, although I think it is probably not the best choice for a library. But from a licensing perspective, xapian is pretty different from SQLite
So there shouldn't be any compatibility issues in practice.