Bleve: full-text search and indexing for Go
blevesearch.com
blevesearch.com
Just recently we merged support a new experimental index scheme called 'scorch'. This new index scheme is designed from the ground up to reduce index size and improve performance. It features:
- a segment based approach, much like Lucene
- vellum FTS for the term dictionary - https://github.com/couchbase/vellum
- roaring bitmaps for the postings lists - https://github.com/RoaringBitmap/roaring
- and compressed chunked integer storage for all the posting details
It's still experimental at this point, but shows considerable indexing speedup, index size reduction, and similar query performance to the old index format used today.
The code for this new index scheme can be found here: https://github.com/blevesearch/bleve/tree/master/index/scorc...
Regarding segments based approach like lucene, means we will need to do segments merging which in my experience quite resource intensive. I don't actually know how blevesearch handles this prior segment based approach.
Merging is required and is indeed somewhat resource intensive. Bleve's current indexing approach has no segments, instead all index data is serialized into a key/value store. This approach allowed us to experiment and plug-in a variety of implementations. Unfortunately, the key/value abstraction limits the way you interact with data, so there are a number of drawbacks. One key gain we get from the segmented approach vs the key/value store approach is that we no longer need to maintain a backindex to handle updates/deletes.
Not yet looking into scorch. So scorch would replace other storage engine like rocksdb and leveldb?
Still very much a work in progress but the core is functional.
There's a question I don't find an answer for by looking at homepage and documentation: does it handle concurrent queries? (I wonder, since I see a store is a single file)
That is, is it something akin sqlite, meant to be used as an embedded engine for standalone applications, or is it fit to be used by a centralized api?
Concurrent indexing is also possible, so long as you can arrange to not put duplicate document ids into batches executing concurrently.
As for usage, it is just a library, so it is typically embedded in a single process (though this can serve multiple clients concurrently).
Distributing the index across multiple nodes is done at the application level. At Couchbase we do this with bleve in a separate project called 'cbft'.
I see, my bad. I saw one binary file named "store" in the directory created by the example code, I thought it would be it.
Thanks for explanation!
SQLite's FTS5 allowed me to load the entire data set without breaking a sweat, but query performance was unexpectedly poor for some combinations of terms with no apparent consistency, and it became unusable due to a temp table I couldn't work out how to avoid when attempting to search in descending primary key order.
It was many months ago now and I have forgotten some of the particulars but thought it might make an interesting jumping-off point for discussion about any ideas anyone might have for indexing 60gb of data into a flat file setup like Bleve or SQLite uses.
Would the new scorch experiment perform better with an index of that size than the boltdb backend?
With scorch the index size is 22MB. Query performance is comparable (and we haven't even gotten to really tuning this yet).
Can you describe how this golang plus elasticsarch hybrid works? Is this is distinct from the link you provided?
I’d love to understand at what scale people are using this engine.
Is that a common thing?
https://github.com/blevesearch/blevesearch.github.io-hugo/is...
You can ignore error codes in go by assigning them to the special variable "_", but, outside of very short toy examples, it is a huge warning sign for terrible code. It's certainly not common to ignore errors in Go.
_ = index.Index(identifier, your_data)
it will also pass https://github.com/kisielk/errcheck check
eventually we can just panic error
if err := index.Index(identifier, your_data); err != nil { panic(err) }
AFAIK, there is no native C/C++ library that comes anywhere near modern lucene. for me the killer feature is that you get analyzers for many different languages out of the box with current lucene/elasticsearch.