RediSearch – Redis Powered Search Engine
oss.redislabs.com
oss.redislabs.com
At the same time, ElasticSearch/Lucene puts considerable effort into analysis at both indexing time and querying time that goes into ranking the search results. RediSearch’s ranking is “user provided”—how does that work exactly? What does that even mean? Ranking is at the heart of information retrieval—of what value is it to allow for 4x more queries to be run on a cluster when the result sets are terrible?
ElasticSearch could be better at scaling up on nodes to handle more query operations on a given cluster. However, if you care about full text search, this won’t do the job.
The benchmarks are quite realistic (5m docs, 25m products) for document search, and the settings reflect real usage. There is no “memory” storage type for ES, it has “mmapfs” which the documentation explicitly says is dangerous and might be removed. The default `hybridfs` already has the ability to use memory mapped files as an optimization. Forcing mmap would actually make it an unrealistic benchmark.
> ElasticSearch could be better at scaling up on nodes to handle more query operations on a given cluster
Based on? They are both running with 5 shards for the benchmark.
I rememeber long ago when SQL Server included full text, we thought that our time fiddling with lucene.net had become to an end, but when we tried it failed miserably because the full text engine and the relational one where like 2 different parts, and if you wanted to make a full text search and order by a numeric field, it would have to make a temporal table with all the results of the full text and order that. Those are the things tha lucene solves so well, that I'm reticent to think that redis search has managed to make them ok at the first try. So, not apples to oranges benchmark, but if you have already redis and the search capabilities that you want to add are fullfiled by redis search, it can be a good product.
Your example doesn't really match the reality of Redis. Modules in Redis can bring their own data types and algorithms and don't need to resort to the same kind of hack you mentioned. The ecosystem inside Redis is designed for modularity and clear, minimalistic interfaces between components.
RediSearch might not be perfect, but any problem will stem from a different set of causes, not because it tries to "emulate" a data model.
The whole point of Redis is not having to emulate data types and algorithms.
To emphasize this:
One thing people might misunderstand about Redis is that Redis extension developers aren't expected to fit their data structures to the needs of a storage-engine. It's not like Cassandra, where your data structures must 'boil down' to key-value pairs; nor is it like an RDBMS, where your data structures must 'boil down' to tuple-sets; nor like a graph DBMS, where your data structures must 'boil down' to EAV triples. Redis data structures aren't "implemented in terms of" any other simpler 'canonical' data structure.
Instead, when you look at something like Redis Streams or the Redis Graph module, the whole complex data structure for each stream/graph is a big opaque in-memory thing dangling off a single key. It doesn't need to be broken down into parts "legible" to Redis. It can just be what it is. "Objects" in the Redis keyspace (the things keys are holding) can't hold pointers to one-another, so the core Redis commands can blindly manipulate them (e.g. deallocate them.) For everything else, you go through the module's commands, which walk the internals of the data-structure as what it is: a plain-old in-memory C struct, defined in your module's header files.
The canonical representation of a data structure in the Redis AOF (WAL log) is just the sequence of commands used to build it; not the data-structure itself. So, as a module developer, to get AOF persistence of your module's types, you don't need to do a thing, other than ensuring that your module's commands are deterministic.
You do need to do a bit of work to get your module's types to serialize into Redis RDB snapshots. But it's completely up to you how to define your types' serializations. Redis just provides an API for writing and reading scalar types from the RDB file stream. How your module uses them to save/load a value of a type is up to you. (And you can just skip this if you like; RDB persistence is used far less often than AOF persistence, so support for RDB persistence it's not even a highly-demanded feature for modules. If you don't bother, then loading your module just disables RDB persistence.)
I believe that the combination of tmpfs + file access — and especially tmpfs + mmap(2) — results in a fast path that almost resembles direct memory access. (If the file or memory-region is opened read-only, the kernel can expose/share the relevant memory pages directly into the process.)
That's completely untrue, it comes with a few scoring algorithms out of the box (https://oss.redislabs.com/redisearch/Scoring/) and results can also be ranked by the value of arbitrary numeric or short text fields. On top of that, there is a plugin API to provide custom scorers written in C/C++.
BTW It's funny though that in the end most of the use of ElasticSearch is for analyzing logs and all the IR scoring tricks are not even used.
Why would I pay to be off on an island very locked into a proprietary search tech? Even if it is a little faster with the redis brand attached to it? It doesn’t seem quite turnkey as Algolia or as deeply featured like Lucidworks Fusion...
Maybe I just don’t get the pitch yet...
It attaches to existing redis hash keys and allows you to add an index to your already existing data for example.
Disclaimer - I'm the original author of RediSearch (started in 2016), before the switch to a non FOSS license. I'm not affiliated with it anymore and haven't followed its development for the past couple of years. But I can tell you that our first users were people who couldn't get ES working for their workloads, or had good redis infrastructure in place already and were able to utilize it for this as well.
The non FOSS licensing mixed with comments about needing an "Enterprise" account confused me TBH...
But TBH it started as a demo dogfooding project for the redis module system - something to demonstrate how powerful the API is, and at the same time test it, find bugs and design problems with it, and gather real developer asks. It became the reference project for the modules API and a lot of the module system's features were designed to accommodate it (most notably async execution of slow commands).
But then a few people started using it (which all of the sudden gave redsiearch itself user asks and testing and all that), requesting features and being happy with it when these features got implemented, and it slowly started to be a thing, with people even contributing code. Then we wrote the distributed version requested by enterprise users, and it became a product, which is now developed by a team. I left it a bit over two years ago.
The backend storage for stream activities is Redis. It's lightweight and fast enough for most use cases.
Sure, not a full on database use case... But it is data that is persisted.
https://redislabs.com/redis-enterprise-cloud/compare-us/
https://redislabs.com/redis-enterprise-cloud/pricing/
thats not particularly expensive, but a dealbreaker for hobby project on which i would want to use it. sadly, there is no non-commercial licence either.
If I buy into the Redis Enterprise Cloud, what are it's preferred usage scenarios, how do I secure the connection properly, etc. etc. etc.
This is just shoving a buy button in my face and I have no idea how to decide if what I'm buying is actually feasible for me, for example from a compliance pov.
You can use RedisSearch on Redis Cloud Essentials with our free plan. We currently support AWS/Mumbai (ap-south-1) and we plan to gradually make it available in other regions as well.
Feel free to take it for a spin and learn about it. AWS doesn't (can't) include it in it's AWS offering so you'll have to buy Redis Enterprise.
EDIT: reh-dih-search, since the "re" is pronounced like in "red" [0] https://redis.io/topics/faq
[1] Squawk - Walkie Talkie for Teams https://www.squawk.to
This is a bummer because high availability is really important for many search uses cases. For e.g. think about e-commerce where search literally prints money.
EDIT: it seems like the open source version supports a read-only replica for failover but my overall thoughts about not crippling/compromising the clustering story in open source version still stands.
While I understand the rationale for this move, unfortunately not having HA in a non-starter.
I had a similar temptation when open-sourcing Typesense (https://github.com/typesense/typesense) and thought long and hard about keeping clustering as part of a closed source commercial edition but eventually decided against it. I understand that commercialising certain features is a necessarily evil and trade-offs must be made. However, I think there are still many avenues to do that without keeping clustering closed source.
Apart from that, I am happy to see the search space heating up with a lot more interesting options.
Disclosure: I have stakes in the open source search eco-system (https://github.com/typesense/typesense) but a genuine Redis fan.
Citation is needed. A counter-point to yours would be how Amazon's search is horrible, because horrible search prints money.
Citation needed as well :) What I actually meant is that for a class of use cases like e-commerce, search is an important feature with a direct impact on revenue. For example, search in e-commerce enables product discovery. So a downtime hits your revenue directly.
Redis Modules created by Redis Labs (e.g. RediSearch, RedisGraph, RedisJSON, RedisML, RedisBloom) are licensed under the Redis Source Available License (RSAL).
Here is their RSAL license: https://redislabs.com/wp-content/uploads/2019/09/redis-sourc...its source is only available so you can look at the implementation if you wish to explore the internals.
> Licensor hereby grants to You a non-exclusive, royalty-free, worldwide, non-transferable license during the term of this Agreement to: ... (b) use the Software, or your Modifications, only as part of Your Application, but not in connection with any Database Product that is distributed or otherwise made available by any third party.
The license is basically to stop AWS using it as part of elasticache (or any other cloud providers)
That said, fuck paying for some bolt-on approximation of a search engine with no real durability. I'll use elastic or solr.
They're free, and have massive communities and user-bases. Issues are surfaced quickly. Compatibility is a priority. And all the other benefits of network effects.
Presumably you're storing your data elsewhere and pumping it into Redis for search
EDIT: I did read the article, though I overlooked the link to comparison with Elasticsearch.
If I were @aphyr, I would say that performance and correctness are competing, so a more performant distributed system is less correct, unless proved otherwise.
What an interesting way to look at performance.
1. What human languages does it support.
2. In these human languages how does stemming and decompounding work in your implementation.
3. how is word importance determined in your index - TF-IDF? other algorithm? Are least important words automatically dropped from queries?
4. Do you have ability to rank on both the stemmed/decompounded query/results and exact matches? So something like raw field access.
5. Can I create my own semantics - I remember seeing a post on here recently where someone had created a search engine (in Rust I think) that was faster than ElasticSearch but from what I could see you couldn't create your own field names so you were stuck searching in title, description, body, creationDate and a couple other fields which really decreases the usefulness.
I mean these are the things that right away spring to mind to ask about when someone tells me they have a new search engine, and when they show me look at my speed benchmarks I'm thinking "what am I supposed to do with this?"
on edit: formatting
on second edit: So I guess as in most things I am interested in how the product actually fulfills what should be its primary functionality, so how does the search engine function as a search engine, I suppose my questions could be answered with quick - our search engine has feature parity with ElasticSearch / Solr where features A, B, and C are concerned - features D and E will be supported in the future.
I also pointed bellow to the specific relevant area in the docs.
> 1. What human languages does it support. > 2. In these human languages how does stemming and decompounding work in your implementation.
https://oss.redislabs.com/redisearch/Stemming/
> 3. how is word importance determined in your index - TF-IDF? other algorithm? Are least important words automatically dropped from queries? >4. Do you have ability to rank on both the stemmed/decompounded query/results and exact matches? So something like raw field access.
I'd like to see actual independent feature and performance comparisons before I come to any actual conclusions.
It seems Redis is moving a bit in the same direction - although not as complex as ES has done it.
Being able to run this inside a Redis instance is a big win, although I suspect few/none of the cloudproviders are willing to pay RedisLabs for the privelige of using the module.
Depending on what you're planning on using it for, make sure to review the documentation, and check the GitHub issues/discussions. For example, I've had to use workarounds to handle some lacking auth/permissions support, but they're currently working on improving it.
Best part by far is the performance. And memory requirements is completely reasonable, depending on the size of the database.
It seems painful to have to write code to reload the search db if it fails.
How long is redis search going to exist and be supported?
If this is delivered as a module, what guarantees do I have that the module interface wont break and leave redis search in a broken state?
Anything that kills redis persistence is also going to corrupt whatever else database you'd have used instead. In fact redis persistence is such a simple model, I'd be surprised if most RDBMs weren't more likely to corrupt data.
I've been running a redis database since 2013 and haven't had a case of lost data once and never even had to restore from a backup.