Unlike most databases the core data structure is the
fastest part of the system. Most of the query time
comes from parsing the REPL protocol and copying data
to/from the network.
I wonder if anyone in the Redis ecosphere has explored a binary client server protocol, something that could be parsed/compiled on the client and then executed without parsing on the server, if the above is really true seems like that might offer even more perf gain than multithreading on the server.I can see cases where a really optimized system could benefit from a binary protocol, but I suspect it'd be a loss for most people.
if (c->argc == 5 && !strcasecmp(c->argv[4]->ptr,"withscores"))
and a quick grep suggests that's a common pattern % grep argv src/*.c | grep -c -e 'str\(case\)*cmp'
482
I guess this means someone would have to tackle creating an intermediate binary format first, rewriting the command handlers to expect that format, and then making client libraries that can produce the format. Perhaps still worth it in the end, but not trivial.[1] https://github.com/antirez/redis/blob/unstable/src/t_zset.c#...
One related choice that Redis makes (or made at the time) is to rely extremely heavily on the malloc implementation, rather than doing work to manage it's memory internally. Even a very trivial, naive free list provided a modest speed-up, for example.
There are a lot of these choices in the code base, largely owing to maintainability concerns (though antirez can surely speak for himself). Given how easy it is for an otherwise uninitiated C programmer such as myself to hack on it, I struggle to disagree with the prioritization. :)
"Unlike most databases the core data structure is the fastest part of the system. Most of the query time comes from parsing the REPL protocol and copying data to/from the network."
SELECT foo FROM Table WHERE key = @mykey;
Then you bind the parameter to whatever you're interested in.In a webserver-like context it's once per query one way or another - the server process is stateless-ish between page loads, so each page load is either a from-scratch connection or a connection taken from a pool, but even if you're pooling you can't use prepared statements in practice (you can't leave a prepared statement on a connection that you return to the pool because you'll eventually exhaust the database server's memory that way, and you'd have to resubmit the prepared statement every time you took a connection out of the pool anyway because there's no way to know whether this connection has run this page already or not).
If you assume a page that's just displaying one database row, which is not the only use case but a common one, then each page load is one query and that query will have to be parsed for each page load, short of doing something like building a global set of all your application's queries and having your connection-pool logic initialise them for each connection.
I'm somewhat surprised at the mechanism you're describing, but now I read the documentation it does seem to be the case. I wonder if a small piece of middle-ware might be sufficient to replicate the behavior I'm describing on a connection pool, and whether that would be desirable.
I am sure it would be possible to provide these guaranties while offering concurrent execution but most likely at the expense of a simple design.
The goal is 100% Redis compatibility so I can’t compromise on atomicity.
https://www.martinfowler.com/articles/lmax.html
I believe other (grown-up/legacy!) exchanges work the same way.
I wonder how much of the direction of concurrency research is driven by the fact that there is much more publishable work to be done in managing concurrency rather than avoiding it!
It also isn't exactly single threaded. It calls fork when it wants to persist data to disk, which is functionally similar to starting up a thread to do disk IO.
Also, it might now have any benefit for you. Imagine a certain key is particularly hot. Having one multithreaded redis process handling access to it might speed things up. Running multiple sharded redis processes won't, since only one of them will have that key.
It would make a little difference if there was a single hot key, but that's a bit unusual. Typically there's some subset of keys that are hot, and you can get them to hash across instances. People also tend to cache those values on the clients, as banging on redis constantly is a waste.
It's pretty common when using a single Redis stream as a lightweight Kafka.
Because it seems like this guarantees that. It mostly parallelizes the command parsing and networking side. The actual core hash table is guarded by a global lock, so you could still get all those single-threaded guarantees.