MeiliSearch: A Minimalist Full-Text Search Engine
tech.marksblogg.com
tech.marksblogg.com
Eventually I ended up on MeiliSearch repo, I fixed an interesting bug, and I must say that the maintainers were super nice across all the process, a couple of months after my contribution they sent hand written letters and a bunch of stickers to all of the project contributors, one of the nicest interactions I ever had on the internet (ironically the first PR that I wrote that got accepted involved one line of CSS, which is field I'm proficient at).
Overall well worth the time spent learning Rust as I feel it makes me a better programmer overall enforcing the thinking about lifetimes, return values, shared data and thread safety.
Their team is fairly responsive to bugs but I had one negative experience when trying to help them fix their instantsearch lib. They were grabbing as many pages as you had set for max pages at once and would re query it on pagination - huge waste of data transfer. They refused to see the problem so I just did a private fork just to get it working but far as I know that’s still a bug.
I need to upgrade the engine itself but looks like they added the ability to upgrade and not lose all the data. That was frustrating but understandable.
Overall I’m very impressed how stable it is
Based on the informations I can read here, I think it comes from the fact that the engine is not able to give an exhaustive finite number of records matching the query for reasons of response time. A finite pagination style (with number of pages) on the client-side is for now a pure work-around.
From what I understand, some of our users try to use MeiliSearch as a primary datastore or expect a classic finite pagination coming from a SQL database env, when we are here to solve search relevancy problems.
Ideally the search results should be relevant enough so that end-users don't have to click on another page selector button, that's why we advocate to integrate a pagination without number selection. Infinite scroll style or prev/next.
Happy to discuss this further with more context!
Thank you for your feedback :)
Im not using as a primary data store.
https://github.com/meilisearch/instant-meilisearch/issues/18
The only argument here that is being made here is that it is 'written in Rust'.
Just use something production ready like Typesense. [0]
[0] https://typesense.org/typesense-vs-algolia-vs-elasticsearch-...
However, today TypeSense indeed has more features than MeiliSearch. After a long time of refactoring the engine's source code, we now have a solid base to welcome new features/improvements, and we hope to evolve quickly to solve many more search use-cases.
For Q3, we plan to add two new features: sort by and geo-search. The geo-search will come out as a first iteration allowing to sort documents around a geographical point and filter documents within a circle. We will also further improve the indexing speed (again yes, because we can do better) and provide two new formats for data indexing (csv and ndjson).
For Q4, we plan to add high availability and solve the multi-tenancy use-case.
That is just a preview of the upcoming features we are already working on.
The end of the year will be rich in evolution for MeiliSearch. We are looking forward to seeing you enjoy using MeiliSearch one day!
So if you see anything that’s wrong in the matrix, I apologize. Please do let me know which items are wrong and I would love to correct them.
I will surely keep a close eye on your developments. Thank you!
"Would it be possible to have a single large highly available index with tenants of wildly different sizes?"
Yes, we can imagine API keys allowing each consumer to have access to a certain number of documents in an index. When querying the index, the internal filters allowing to select the documents accessible by this API key are automatically inferred and added to the initial query of the consumer. In short, it's a bit like a WHERE clause that would be fixed at each query.
I'm not sure if it answer your question, so don't hesitate to reach me!
I'm available to discuss your use cases from our community slack or by email :)
Finally companies which work in a far more human way.
And I personally stopped using them after a really bad experience I had with their "developers". They don't really care about you and it shows, also, they were kind of rude when I reported some bugs to them.
I moved to typesense and it's a whole different world, their creators truly enjoy that you're using their product; same thing with sonic, Valerian is the kind of hacker you'd want as a friend, super talented, super easy going, you could ask a completely dumb question on their GH and he takes the time to explain things to you at length. I know its open source, I know I didn't pay a dime, but for me, that kind of attitude makes it or break it. Plus, you actually get a superior product.
What? Are you me? Haha.
Check your email, hermano!
[0]: https://typesense.org/blog/the-unreasonable-effectiveness-of...
One of my favorite parts of working on Typesense is the opportunity to interact with so many developers from around the world, getting to know about the product and domain they are working on, their tech stacks and how Typesense fits into their world. I find these interactions helpful in enriching my own world view and helps me build valuable context as we design new features. I’ve sometimes been blown away by how the foundational construct of a fast and distributed search engine, is being used for use cases I could not have even imagined!
Moreover, we are certainly, on some features, maybe a little bit late. Delay that we will more than compensate before the end of the year. Our priority until now has been to offer a robust search engine accessible to all. For us, the developer experience is really important, whether it is in the use of the API or in the communication with the community.
We will continue to try to do our best for the community. If you want to help us to improve, I would be happy to take your feedback.
Keep up the good work.
I’ve commented on both Stripe and DigitalOcean’s terrible support, under different accounts, and had some XO give the “we’re so sorry, email me directly and I’ll personally look into it, our customers are really important to us” tripe.
There is no good work here. Only platitudes.
Across the assortment of Meilisearch repositories, I've raised two PRs (one accepted, one rejected), five issues, one feature request and pinged one issue for an update.
Every single time the Meilisearch team has been responsive, communicative and generally a delight to interact with - there are very few projects I would consider better.
Just thought I'd throw in my experience.
Can you please DM me your concerns @rrjanbiah? I'll try to coordinate with the Meili team and sort them out.
Wow; how on earth can this blow up to 35 MB? For comparison: the Crossline stand-allone exe (http://software.rochus-keller.info/CrossLine_win32.zip) with built-in https://github.com/rochus-keller/Fts and Sqlite (all written in C/C++) is less than 7 MB. Where do the other ~30 MB come from?
It is great to put a concept in place. For more advanced use (mainly index and search features) I was also evaluating TypeSense which didn't win me over as a product. I have not tried Algolia because of perception that it is heavier and paid from get go.
Arguably I'm not the biggest fan of ElasticSearch, it's a way too complex to manage and interact with, if you just need to add search to a product. However, ElasticSearch i also much more than just a search engine. I would never use Bleve or Sphinx as a primary data store, but ElasticSearch is a perfectly good document database.
I recently asked about this and people replied that it wasn't fit for this purpose
However, that's not really what it was made for. Especially early on when you're planning out your schema and such, dropping and re-indexing your documents is a really simple task. If the index itself is your primary document store, what are you indexing from? Would you have a DBMS or file system as your secondary store in that case? That just seems so awkward and backwards.
Keep the square pegs in the square holes and use Elastic (and the alternatives discussed in this thread) as a search index.
https://correlate.meetglimpse.com/
If you're doing some test products and just want to have a search that is easier to setup than ES. Meilisearch is a great alternative.
Insertion times grows linear with index size, up to tens of milliseconds with an index of couple 100k documents.
Go library is very un-go, with not all the options exposed. And had a couple of breaking changes without upgrading major versions.
Other then that, the search part works really well
MeiliSearch doesn’t strip HTML tags and i had to do that manually before adding posts to index
However, after thinking about it more, I wrote up this issue[0] with some ideas and thoughts so I could implement it as PR or work around it.
I ended up working around it, because that makes most sense: separation of concerns: meilisearch should indeed not get involved in stripping or fixing HTML as that i) ties Meili to HTML, ii) requires configuration and complexity to allow control and iii) adds features that become security-critical.
Indeed, my solution is to sanitize, clean and strip HTML before sending into the index.
I found the default order of results a bit off. Near-matches were positioned over exact matches.
I'm looking for a fulltext typo-tolerant search tool that integrates well Hasura+PG.
No nice integrations like hasura-backend-plus and combines hasura with minio/s3 and authentication service.
Use PG's built in full text search capabilities:
https://hasura.io/blog/full-text-search-with-hasura-graphql-...
https://www.lateral.io/resources-blog/full-text-search-in-mi...
Extend those capabilities with pggroonga:
It's ridiculously easy to use and has faceted search for my needs. However, there are some limitations so I have to use it in combination with redis, but the developers have a roadmap to fix these problems.
Synchronising with MeiliSearch is a bit of an effort because of the following limitations:
* When filtering by facet, it doesn't provide count for disjunctive facets
* No sort by
* No where clause (less than 50 for example)
To overcome these problems, I rebuild some parts of the database in redis, use code for filtering and query MeiliSearch multiple times for different facet counts.Both redis and MeiliSearch are ridiculously fast so the performance loss is negligible, but it makes my code quite complex. As soon as the developers add these missing features, I want to simplify my code and only use redis for query caching. Typesense had some of these limitations too, but I'm not sure if that's still the case.
Concerning the disjunctive count of the facets, we are thinking about it. It is feasible on the client side by making several requests but we are aware that is it not ideal at all from a developer experience point of view. We are still thinking about the best way to solve that case in one of our future iterations!
The sort feature is coming in v0.22 (string and numeric fields) you will be able to easily configure the balance between exhaustivity and relevancy at index level through the positioning of the ranking rules.
I'm not sure I understand the where clause point so I'd love to hear more details!
Thanks for using us and giving us this kind of feedback :)
By where clause I mean as in SQL. For example, select results where cost <= 50.
[0]: https://www.elastic.co/pricing/faq/licensing [1]: https://news.ycombinator.com/item?id=28110610
Looks like it was designed with the user in mind. Telemetry by default.
https://typesense.org/typesense-vs-algolia-vs-elasticsearch-... (yes it's hosted by Typesense)
https://github.com/typesense/typesense/issues/122
https://github.com/typesense/typesense/issues/95Oh
The workaround they recommend is to duplicate your index with all those characters removed and then strip out those characters from your search queries :/
It's an information that is typically missing yet very important!
The real advantage of LMDB is that it is a BTree, key-values are ordered and do not need any computing when retrieved which is not the case of a LSM-Tree key-value store like RocksDB that needs to merge/compact pages of key-values pairs before being able to return it too you. Wasting CPU when the search engine must use its CPU to do union/intersection…
Another advantage of LMDB is that it returns a view into the DB itself of the entries, RocksDB can’t as it must do operations on the entries before returning them to the library user, for example: decompressing or compacting the values.
In the case of Typesense, RocksDB is not even a top-10 contributor to the overall latency involved in serving the result. In any case, it would be good to clarify a few things:
> As we don’t use RocksDB but LMDB, we use a lot less real memory than key-value stores that uses a user-side cache system.
Typesense stores only the raw-data in RocksDB. All indexing data structures for filtering, faceting etc. are compact in-memory data structures stored outside. The only fixed memory cost from RocksDB is an in-memory table that is used to buffer writes (see the next point) before being flushed to disk. In practice, this is a trivial percentage of memory used when compared to other data structures.
> LSM-Tree key-value store like RocksDB that needs to merge/compact pages of key-values pairs before being able to return it too you
This happens in-memory and is flushed to the disk in batches. Merging of on-disk SST files happens in the background with no real impact on reads. The advantage of this approach though is that it gets you really good batched write throughput [0] (the above caveat on the difficulty of benchmarking applies).
In summary, like all systems, choosing a storage system involves many trade-offs and what really matters is what works best for your architecture.
RocksDB doesn't support transaction but views in the database, which means that if you are indexing, writing into the database and that any event makes your program to stop unexpectedly, you can't just start your program and use the data like this as it could be corrupted.
This is why at MeiliSearch we prefer using LMDB, even in case of an unexpected crash, a reboot is instant and valid, you just need to restart the indexing you were previously doing and can serve requests to the users with the previous version of the database.
Also as you can see [0], the benchmarks between LMDB and RocksDB is very clear. I understand that it is maybe not reading the database that takes time on your side but it is on our side, combined with the set operations between sets of internal documents ids.
All the best with MeiliSearch!