Manticore 6.0.0 – a faster alternative to Elasticsearch in C++
manticoresearch.com
manticoresearch.com
It works best when you have a SQL data store you want to index against, but with the real time index you can treat it more like elastic and other searches. However for that first use case of SQL, I don’t know of anything else that comes close to being as easy to use.
Simply point it at your database, give it a query to pull what you want to index and you are done. I suspect that this covers about 90% of use cases out there.
If you need more than what the DB native indexing is giving you give it a try.
This sounds like a bit of a killer use case but even explicitly searching the documentation I can't find more than tease level information about it. They seem to be hyper focused on just presenting it as a replacement for ElasticSearch.
Probably it's because they think it's where money floats, seeing ES as a train to push them forward as well.
Not long time ago, ~ 2 years of so, had a case where dev team (okay, it was just 2 persons of developers) asked to setup ES to offload some searches from DB (mysql). I've asked why not using Sphinx/Manticore and one of the reasons they said they want ES is cuz in Laravel [php framework] they have some nice libs to work with it, while for Manticore it seemed to be more work for them. Quite shocking reasons from my POV, but real story.
How did they know in the previous 5 major versions which part of the product to improve?
Actively telling a project how to improve vs passively observing how the system is used in the aggregate: that is a choice each person should be offered to engage with the project, along with the right to lurk in peace or fix their problems their own ways.
The data itself ought to be public, and if it can’t be then it shouldn’t be gathered. The insights from that data should come from the community, not just project leads. That data’s relevance to product development and design should be annotated on feature tickets.
I am good with having a relationship with a project that involves sharing my behaviors within the system, others are not. As a user I do I want to be able to see if I’m way off the beaten path or using a popular method when I’m deciding if I should debug, research, or work around a problem. As a maintainer team, your time is precious and the BS is already thick, so I feel providing that visibility is the least I can do to contribute to your work.
I feel very differently about this with respect to for-profits and my PII tied to behavioral records.
If end-users wanted "telemetry", then they would have asked for it in previous versions.
Even once telemetry is added, there is still no reason to enable it by default. End-users that want/need to send data to the company, e.g., to back up their claims about deficiencies in the software, can easily enable it.
Show us the contractual restrictions that limit how the data sent by the end-user can be used. Show us the guarantee or enforceable promise that this data collection will result in improvements that end-users want (versus only benefitting the company in some undisclosed way(s)).
Why would anyone want to send behavioural data to a company with no enforceable promise of a benefit and no way to monitor how the data is used.
Manticore Search: Elasticsearch Alternative - https://news.ycombinator.com/item?id=32261618 - July 2022 (69 comments)
I found an article from 2022 that did a compare/contrast but I wanted a feature by feature breakdown.
If the data that you want to search is entirely contained in a SQL database, it's an uncomplicated and powerful solution, definitely check it out. If not, Manticore may still be a nice solution for you, but I can't speak to that.
it's been rock solid for all those years.
Manticore is different in terms of this, especially the Manticore columnar storage which doesn't require a significant portion of the data set to be stored in memory. This allows, for example, for a 1TB data set to be served on a standard server with some 32GB of RAM.
Mantiscore is an opensource fork of Sphinx Search, which released its first version in 2001. The fork started after the latter went from opensource to proprietary, at the end of 2017. The engine is stable and battle-tested. IIRC, Craiglist uses Sphinx.
I have been following these two libraries (Manticore and Meilisearch) very closely. Their simplicity, portability and performance gains over Elasticsearch are impressive.
Since two days ago, I am creating Python bindings for the core search engine of each of these two libraries, starting with https://github.com/AlexAltea/milli-py. Getting extreme performance, but as an embedded/self-contained package (basically same goals as SQLite).
Grapana Loki advertises lower resource requirement, but it's just a disk storage system. Any query will read everyrhing from disk.
The Elasticsearch has big RAM requirements if you create a lot of indexes of course. You can't have something more quick than indexes, and you can't have lower resource requirements without having fewer indexes.
That's what I'm asking, actually. Isn't Loki's proposition that it only indexes the tags and time interval? Do you mean that even filtering by that there's still a lot of data to go through?
Because it seems like you're saying it always fetches everything from disk.
Yes
> Because it seems like you're saying it always fetches everything from disk.
If you specify a tag, like environment, it will not read the disk for data from other environments. But the tags like environment/host/timeframe are not enough if you want to query for something like error/exception/sessionid, and you might have to wait minutes/hours for a query which covers a lot of data.
What can be a cool feature, it's auto backup to S3, or load from S3.
You can get around that by having the search happen in a separate process or something, maybe. But this is a huge issue for something that one might want to embed.
the GPL bleed only happens if you distribute your application, meaning to sell or give away binary packages for customers to install. if your product is a hosted api that you do not distribute, you do not invoke that clause.
also, a lot of open source projects handle this by having things like the core engine licensed on a copy-left friendly license (GPL,AGPL). however, the language connectors and bindings are licensed under the slightly less restrictive apache license. unless you are offering a saas service of the product itself, it is more likely you are actually interacting with the connectors anyways. mongodb is a classic example of this model.
The point is that I can use SQLite, Tantivy, RocksDB, ... in my app no problem. I can make it open core, I can make it AGPL, BSD, MIT, not problem. Because those things are meant to be embedded. But I almost definitely can't use Xapian.
Let's be honest, if I want a search solution for use in my SaaS, I will grab Elasticsearch or an equivalent, I have no need for a library. It seems to me that the only use case where Xapian could really shine is crippled by their license. That is a shame.
* License preference: Some people prefer true open-source licenses as opposed to the license that Elasticsearch has switched to.
* Performance and resource consumption: For some, performance and resource consumption are significant factors in their choice of a search engine.
* SQL vs JSON DSL: Some people prefer using SQL over Elasticsearch's JSON domain-specific language.
* Maintenance: Some believe that maintaining Elasticsearch can become challenging when the data collection becomes large enough.
That's what I've heard from those who preferred Manticore over Elasticsearch.
Those people can also use OpenSearch, which is a recent fork of ElasticSearch (by Amazon) that is using the Apache 2.0 license.
Elasticsearch is a pain to tune and partition, and the JVM brings a whole set of operational issues but what's the point of better read/write performance when the actual search performance is worse?
I guess this makes sense for use cases where you care more about speed than the quality of results.
[1] https://docs.google.com/spreadsheets/d/1_ZyYkPJ_K0st9FJBrjbZ...
This is a very strong claim, and without strong arguments, it's a ridiculous claim.
Looking at the docs I could only see _create and _doc but not _bulk endpoint support. How will that work with Logstash and Filebeat?
lnx might be similar, I'm not sure. It's very new and I had a bad experience trying it out.
https://github.com/manticoresoftware/manticoresearch/
https://db-benchmarks.com/test-taxi/#manticore-search-vs-ela...
However, with good enough algorithms and judicious coding and memory management, the possibility exists.
Plus it doesn't need to be Java xor C++, JNI exists for a reason (now Panama).
Languages such as Java or PHP make you lazy and you end up using the string variable type a lot. It is extremely inefficient.