Sonic: Fast, lightweight and schema-less search backend
github.com
github.com
Even if it offers only a fraction of the features offered by ES, that may be fair enough for at least half of the use-cases out there.
Sonic could have really had a strong selling point: "Use an ES-alternative that works fine in most of the real-world applications, but it's written in Rust and it only takes a fraction of the memory footprint required by ES, and it shouldn't require you to change your application code".
Instead, they are proposing yet another search protocol, that developers have to learn and adopt. That definitely increases the adoption barriers.
Go client: https://github.com/elastic/go-elasticsearch/blob/3985f2a1554...
Python client: https://github.com/elastic/elasticsearch-py/commit/e72aa3e24...
my guess is because it's "X-Elastic-Product" you'd be falsely "saying" your product is made by Elastic, the company so they could sue over it.
https://stackoverflow.com/questions/1114254/why-do-all-brows...
Such a concern seems utterly ridiculous.
But yes, Sonic could replace lots of use cases.
Right now I've got the site going on just Postgres FTS + trigram and it's pretty darn fast, looks like I need to test sonic too.
Going to burn some midnight oil (in my timezone, anyway) and get it out -- though sonic isn't implemented yet!
Anyway to make this comment useful to people, here's my short list of engines that I want to run in parallel:
- MeiliSearch (https://github.com/meilisearch/MeiliSearch)
- TypeSense (https://github.com/typesense/typesense)
- Lyra (https://github.com/LyraSearch/lyra)
- OpenSearch (https://github.com/opensearch-project/OpenSearch)
- ZincSearch (https://github.com/prabhatsharma/zinc)
- Sonic (https://github.com/valeriansaliou/sonic)
There isn't enough out there comparing all these for the simple typical fuzzy search/search box usecase, so I'm adapting a little podcast search site I made to try and use all of these at the same time. So far only Postgres though, will try and add Meilisearch today and post it!
Like other people are pointing out, most of these engines won't have all the features of ES (or more accurately Lucene) but I am pretty convinced that most of the time it doesn't actually matter and if someone is searching on your site excessively maybe there's a problem with your UX (unless you're a search engine or repository of information).
[0]: https://supabase.com/blog/postgres-full-text-search-vs-the-r...
I don't understand this comment. Why would you search something that *isn't*, in some senses, a repository of information? I would say almost every website needs to have search in some sense, and it's *because* sites function as a repository of information that they need this search. Think about e.g. Stripe's documentation, or Github's repository / code search. HN is also another great example—I search for stories or comments all the time to try and remember something I read about recently or heard about last week, but couldn't quite remember. I'm hard-pressed to think of a web site I use regularly that *shouldn't* have full-text search, if I'm being honest.
Just gooling site:foo.com/baz <query> almost always produces better results.
The scale of a documentation site is a very different problem -- you can brute force it in ways that you can't at larger scales.
I agree that HN would be a case of the large repository, but even then what most people want out of HN search is pretty simple/basic keyword search. I think a decent non-frustrating HN search feature could be very basic and get by without most of the advanced features/rabbit holes available in search.
Basically I think most apps fall into the lighter search use case -- command palettes, search inside of apps with a small scale of information, etc.
My comment wasn't that apps shouldn't have full text search -- it was that most that have full text search don't need complex full text search with all the bells and whistles that lucene and other serious search engines provide. These up-and-comers might be enough for a bunch of apps for which search is not the main feature.
Are you aware of any that can be used client side like Lyra and supports faceted search?
I've been looking for a solution and cannot find it, even an algorithm and/or a data structure can be helpful. I attempted coming up with a solution myself but ended up with frustration when it came to making the facets dynamic and update as other filters are applied.
I read a couple of papers and one stood out [0], which introduces category theory as a solution to faceted filtering. I understood it in theory and it was still does not seem straight forward to implement but I haven't attempted yet.
0. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5145200/#!po=28...
https://lunrjs.com/docs/index.html
There are some others but I can't find them at this moment -- a bunch of the other projects I find are somewhat abandoned, lunr is actually on my list of things to use (because it makes the most sense to just ship a pre-built index with the first like... 5 letters maybe of typeahead, no matter how fast the backend is)
I have for many years now a small search engine project in my free-time pipeline, but I'm before crawling even and I intend to sit for searching part after some of that.
Eventually all of these projects will be highlighted on Awesome F/OSS (https://awsmfoss.com), but for now I'm just going to dump my bookmarks here for other people, since I'm leaving awesome projects out:
Search Engines
AWS OpenSearch https://github.com/opensearch-project/OpenSearch
https://github.com/opensearch-project/OpenSearch-Dashboards
https://github.com/opensearch-project/perftop
https://github.com/go-ego/riot
https://groonga.org/ https://github.com/groonga/groonga
https://github.com/meilisearch/MeiliSearch
https://github.com/mosuka/bayard
https://github.com/nezaboodka/nevod
https://github.com/searx/searx
https://github.com/stryku/okon
https://github.com/toshi-search/Toshi
https://github.com/typesense/typesense
https://github.com/valeriansaliou/sonic
Algolia
https://github.com/marconi1992/algolite
https://github.com/quickwit-inc/quickwit
https://github.com/prabhatsharma/zinc
phalanx https://github.com/blugelabs/bluge https://github.com/mosuka/phalanx https://github.com/mosuka/blast
ManticoreSearch
https://github.com/manticoresoftware/manticoresearch
https://github.com/manticoresoftware/docker
https://manticoresearch.com/blog/manticore-alternative-to-el...
https://manual.manticoresearch.com/Introduction
https://forum.manticoresearch.com/t/manticore-search-cheatsh...
https://forum.manticoresearch.com/
Whoosh https://whoosh.readthedocs.io/en/latest/
https://pypi.org/project/Whoosh/
lyra https://github.com/nearform/lyra
https://nearform.github.io/lyra/
https://github.com/LyraSearch/lyra
flexsearch
https://github.com/nextapps-de/flexsearch#performance-benchm...
Lucene
https://github.com/apache/lucene
ZincSearch
Solr
https://solr.apache.org/operator/
https://solr.apache.org/guide/solr/latest/getting-started/so...
https://github.com/apache/solr
https://solr.apache.org/guide/solr/latest/deployment-guide/s...
Konnu https://gitlab.com/shadowislord/konnu
Quickwit QuickWit + Clickhouse
https://clickhouse.com/docs/en/guides/developer/full-text-se...
https://clickhouse.com/docs/en/sql-reference/functions/strin...
There is no way I can get to running all of these (this project was supposed to be quick!!), but I will run the ones I noted earlier, and probably manticore too since it was high on my list since it's quite polished looking.
The list has value on its own, especially if you maintain it.
Waiting for meilisearch to ingest documents right now and the Show HN is going up.
https://news.ycombinator.com/item?id=33321268
Maybe I should have used their batch thing instead.
It underlies a concordancer GUI called Smyrna: https://github.com/nathell/smyrna, https://smyrna.danieljanus.pl
I haven't touched it in six years, other than a few small changes. But I do plan on revisiting it when time permits.
Does this mean that it only ever finds at most N documents per word? Even searches for "A and B" would probably not find everything, even if less than N documents contain A and B, because they might have been removed with the sliding window already for A or B alone. Is that correct?
Every time you think it’s somehow magic, someone has to dump a bucket of cold water over your head.
This is surely an impressive engineering feat, but hardly a replacement for the myriad of query possibilities Elasticsearch offers.
"Sonic can be used as a simple alternative to super-heavy and full-featured search backends such as Elasticsearch in some use-cases."
Seem pretty up-front about it, and doesn't claim to be a full-featured alternative.
I found myself last week reimplementing 10% of RoaringBitmap's functionality as a homebrew replacement, because doing so was 500% faster. Not that RB isn't great, but it's designed for a general problem space, and not my particular problem.
> Sonic can be used as a simple alternative to super-heavy and full-featured search backends such as Elasticsearch in some use-cases.
Having a lightweight search engine is fine, but calling it an alternative to Elasticsearch is not doing either justice.
Vector queries aren't niche, Elastic however only tacked on a proper (non HNSW) implementation in the last year and a half. Geospatial isn't niche, anyone working with location data will work with those queries. TF-IDF is a basic ranking algo / signal.
Maybe Elasticsearch is good for you because they have all their features in aggregate. But I can name a tool that focuses specifically on each area and query type and is better for that specific subset of functionality.
So my point still stands, if all you need are specific features Elastic is too much. You need all of it and that's fine too.
Please do name them, because I for one would like to never run ElasticSearch again for faceted, full-text and specialised search.
Which is perfectly fine. A lot of tools become so general and bloated, that there are large groups that would be fine with many different 10% subsets of their features...
Kind of like how I don't need MS Word or OpenOffice Write, any simple text editing program with a few basic features (like printing, bold/italics, and word count) will do for my needs...
The only exceptions could be small single feature utilities.
The friction is introduced where it’s not made crystal clear how it’s similar, and which concept are different or missing altogether. Then it will cause unmet expectations.
Perhaps it’s better to say “inspired by …” or “similar to …” to make a more precise statement.
Tyoesense is probably the most compete competitor in that regard: https://typesense.org/
Other alternatives here: https://gist.github.com/manigandham/58320ddb24fed654b57b4ba2...
With Elastic Search many of the features, security being one, are locked away behind commercial licenses. With Meili it seems they are, for the time being anyway, going with a proper open source version. I understand Elastic needs to earn money, and I get their licensing model to accomplish this. But Meili will probably steal away a good portion of customers interested in self hosting their search solution.
- written in rust or maybe just C - extremely lightweight and high performance - single small binary that runs anywhere - designed to run in Kubernetes from the ground up - scales dynamically up/down - zero downtime upgrades - rigorous security built into the core offering - fully open source - wire compatibility with ES
I hope that ES themselves do this. There are pretty significant barriers to creating a serious competitor to ES (unlike something like MongoDB for example which seems to have a very limited role in the future).
The truth is that ElasticSearch/Solr/Lucene is orders of magnitude more complex and powerful than these "alternatives". All this is mostly fine as long everyone is on the same page regarding the expectations.
Most people don't need ElasticSearch for their use cases on the surface, but I feel they expect top-notch mind-reading results and that requires something like ElasticSearch and someone who knows the field.
Having said all of that, Meilisearch and this are quite fine.
I migrated from typesense to Meilisearch on a project after I found it had much better search accuracy. I can't exactly explain why, but overall Meilisearch results feel more relevant by default.
[1] https://github.com/beir-cellar/beir [2] https://docs.google.com/spreadsheets/d/1_ZyYkPJ_K0st9FJBrjbZ...
I’d love to take a closer look.
I migrated in April 2021 (latest version of typesense & meilisearch at that time).
I don't have a public dataset has it was a fairly large ecommerce catalog with close to ~500k entries. And again, it was just my own perception which is hard to define. I just found that Typesense was a bit off compared to Meilisearch on search accuracy, and of course could totally be different today with a more recent release.
We're now at v0.24.rc, and we've iterated quite a lot on improving relevancy since then, as more users shared their datasets with us and gave us feedback over the last 1.5 years.
If you get a chance to try out Typesense again in the future, I'd love to hear how relevance feels with the latest version, out of the box for your dataset.
An example dataset is BEIR(BEnchmarking Information Retrieval), published in NIPS 2021: https://github.com/beir-cellar/beir
If you utter the phrase "I just want search" then it really is a matter of just using one of these lightweight projects and libs because your needs are simple.
It won't show you which one is "best" (for a given value of "best"), just one that looks most similar to the input.
Trying to index anything that can contain any trace of SEO would be doomed to failure, it also won't tell you which of the sites got linked the most, and million other things other web search engines do to give good results.
In "just put a documents in DB and search them" it is barely enough to look thru corporate knowledge database and it still won't get nuances like "this page is linked from 20 other pages, maybe it should be higher?"
* Once imported, the search index weights 20MB (KV) + 1.4MB (FST) on disk;
This is almost unbelievably succinct! If you encode the document features into 8 bits per document, and thus completely forego the need to store the document ID by indexing them implicitly, that alone is 1 MB.
Getting meaningful search out of on average 21 bytes per document seriously impressive.
[For reference, this sentence is 42 bytes.]
> Sonic only keeps the N most recently pushed results for a given word, in a sliding window way (the sliding window width can be configured)
Default window looks like 1k documents. I read this as saying that super common words are basically dropped from the index (only 1k out of many thousands of docs retained), but I don’t know enough about the internals to be sure. Not sure if this actually hurts search results in practice, seems like an ok trade off for help docs at least.
Looking at it from a practical example such as log search (almost everyone I know has used kibana/logstash/elasticsearch at some point): you'd be able to search for things like tracingId/requestId but adding more filters such as logLevel, requestType or serviceName would be impossible
It has it's niche, but calling it an elasticsearch alternative really is a stretch
Apart from Sonic, I also found Tantivy [1] and Meilisearch [2]... all delightfully made in Rust. My favorite, and the closest one to ElasticSearch (for its features) is probably Tantivy.
I'd recommend anyone to check up this three projects and choose on what best fits your needs... it's awesome to see that more projects are becoming available by the day!
You can have a look at lnx (https://lnx.rs/) that is based on tantivy and is performing quite well. It's not yet distributed but the author Chillfish8 has some thoughts about how to do it.
Sonic here only returns document identifiers so you will never be able to get document information back. This is very useful though if all you want to do is index text data and then get the stored information from another data store.
Quickwit supports Elasticsearch compatible bulk indexing API.
In many cases that is what you want because you have the data in a database and don't want to duplicate it in Elastisearch.
Why would you want that anyway? Always thought it was silly to duplicate all your data which will be stored in a real database anyway
Also, now that I think of it, typically logs/structured data is stored only in ES.
If you discard many potential hits, why not use /dev/null as the search engine?
They let you configure the number of expected results to cache for a given query, the number of cache results are configurable based on your use-case for the results (e.g. if your website only lists 100 results, don't store beyond that).
If more results than that for a given query are returned then they disregard additional results since you told it you won't make use of them. In essence, they're saving you from caching results that you'll never consume.
How you got from this to "just use /dev/null" is a mystery to me. It has to be a misread or misunderstanding.
SQLite has stemming only for english out-of-the-box, but I find it quite a need for a good ES drop in replacement.
My two cents
That service worked beautifully. Results were returned in 10-20ms and we only ever made software updates to handle the occasional CVE. It did, however, take quite a bit of fiddling initially to get the query results to match the user expectations. For example, weighting first vs last vs full name.
Does anyone know of any alternatives that support this use case?
https://www.elastic.co/guide/en/elasticsearch/reference/mast...
[1] https://play.manticoresearch.com/pq [2] https://manual.manticoresearch.com/Creating_an_index/Local_i...
He built two distributed search services:
- https://github.com/mosuka/phalanx, written in Go.
- https://github.com/mosuka/bayard, written in Rust.
Sonic is good. Typesense is probably what most are looking for as more of an Algolia-like setup: https://typesense.org/
By the way if you are looking for lightweight "alternative" for ElasticSearch you might look at sphinx search engine (although it doesn't has as much features as ES has and it has became closed-source since 3.0 version).
Manticore Search [1] forked from the latest open source version and has been continually improved for more than 5 years.
[1] https://manticoresearch.com/
> although it doesn't has as much features as ES has
Manticore unlike Sphinx is much closer to Elasticsearch in terms of features set.
There’s not much mention of that. I’m always on the lookout for something lightweight that improves on PostgreSQL full text.
Sonic: Fast, lightweight and schemaless search back end in Rust - https://news.ycombinator.com/item?id=19471471 - March 2019 (39 comments)
There are already millions of solution to have entries in an inverted index that you can query within ms, none of them have the power of ES in term of features, scaling and HA.
Looks like it rebuilds the whole index periodically and that's very processor intensive. The delete will be reflected after a rebuild.
[1] https://manticoresearch.com/blog/manticore-alternative-to-el...
I'd like to avoid compiling LLVM from source if I can
One thing is to find what you search, but the other is not to find what you aren't allowed to see.
Comparison to elasticsearch: https://docs.meilisearch.com/learn/what_is_meilisearch/compa...
Github: https://github.com/meilisearch/meilisearch
Website: https://www.meilisearch.com/
(disclaimer: I'm one of the cofounder)