Build an Elixir Redis Server that’s faster than HTTP
docs.statetrace.com
docs.statetrace.com
But if you want better performance, you use active mode. In active mode, the runtime is receiving data on your behalf, as fast as possible, and sending it to the socket's owning process (think a goroutine) just as fast. Data is often waiting there for you, not just already in user space, but in your Elixir process's mailbox. (Also, this doesn't block your process the way recv/2 does, so you could handle other messages to this process.)
You could imagine doing something similar with 2 goroutines and a channel. Where 1 goroutine is constantly recv from the socket and writing to a buffered channel, and the other is processing the data.
One problem with active mode, and to some degree how messages work in Elixir in general, is that there's no back pressure. Messages can accumulate much faster than you're able to process them. So, instead of active mode, you can use "once" mode or {active, N} mode (once being like {active, 1}. In these modes you get N messages sent to you, at which point it turns into passive mode so you can recv manually. You can put the socket into active/once mode at any point in the future.
Elixir sockets can be safely shared between processes, which might not be obvious to an Elixir programmer.
To the best of my knowledge, Elixir will use writev where possible, so it's iolist friendly and can be extremely efficient.
Binary pattern is a productivity boost when it comes to networking work.
Every Elixir socket is associated with a "controlling_process". This creates a link between the socket and the process. The linked article uses it. You generally don't want this to be the acceptor loop, since if the acceptor loop crashed, it would close every socket. Fun fact, I believe earlier versions had bugs / race conditions with respect to changing the controlling_process while data was incoming. This has since been fixed, and things "just work" like you expect, but I can only imagine that it involved some coding gymnastics to fix.
Since the Elixir sockets are more abstracted, there's more knobs you can turn and tweak, e.g., you can tweak the buffers that Erlang is using to read/write. Or it has built-in parsers for common formats, including simple things like automatically being able to send/receive 1, 2, or 4 byte length-prefixed messages.
Elixir has two socket APIs. The traditional gen_tcp, and a new one, socket, which is meant to be "as close as possible to the OS level socket interface." I haven't tried the new one yet.
Might be interesting. Amongst other things, they experimented with a patch to Cowboy to use {active, 100} instead of the default {active, 1} and got some nice wins.
I feel like there's this distressing trend where web programmers export their practices everywhere, and over time, those become the new accepted "primitives" instead of the actual primitives. What used to be a socket is now apparently a whole black-box protocol, where swapping one protocol with another to get speed gains is now noteworthy, and what used to be just FFI or a shared library is now a whole client/server architecture. This is basically the intellectual lineage of stuff like "language servers" where every keystroke your editor heap allocates a JSON tree, serializes it, initiates an HTTP connection, sends it over, reads the response, deserializes it into more heap allocated nodes, etc.... when the actual stated goal (making code intelligence consumable by any client) could be achieved using much less.
Standby as I follow up on this with my own post, "Build a TCP server that's faster than the Redis protocol."
May not be for this instance, but if people are gravitating towards a bad tooling alternative must be worse by some important parameter.
My Mac mini runs at three whole billion clock cycles per second. If we’re going to depend on hardware to make our argument, we’re all basically running supercomputers. And yet everything is slow, and my tools lag for no reason, the web in general is slow. Which is all fine, but I thought computers were supposed to be fast? Somehow we’ve arrived at a point where on the one hand, computers are supposed to be fast enough to support these huge skyscrapers of abstraction, but on the other hand… the result is mind blowingly slow.
I could maybe see some really specific use-cases for this but for probably 90% of cases implementing a distributed process API should be able to go over pubsub fine shouldn’t it? Or does pubsub have some sort of massive overhead?
EDIT: Just because it’s tangential to the topic at hand, and since I’m up here at the top, I also wanna throw out a nod to libraries like `nng` and `nanomsg` which are spiritual successors to ZeroMQ and they have functionality like brokerless RPC/pubsub built in as messaging models. I don’t see tools like this talked about a lot in this space cuz systems software isn’t the sexiest but if you need to embed a small and lightweight messaging endpoint in your backend stuff then look at those as well. No horse in the race, just like sharing useful tools with people.
I often build TPC based "protocols" that are just newline delimited json in each direction, it's a nice middle ground between using an HTTP POST and something like gRPC.
The only use case that really jumps out is when you don’t want a broker because you don’t want a single point of failure but now you’re embedding a Redis server implementation in every one of the services in your mesh and I’m not convinced that’s much better but I can see where it might be helpful.
This is a super cool write-up and I’m not saying anything negative about what the author did. I like it a lot. I’m just asking from a technical and curiosity perspective what the advantages are to this over using the stuff Redis Server already provides which can do the same thing.
I haven’t looked at the exact performance characteristics but it would be fun! You would have built in load balancing!
You can do a similar thing with Elixir and your own protocol (or the Redis protocol).
This pattern is very common in banking but require something like RabbitMQ or Azure service bus …
> The only use case that really jumps out is when you don’t want a broker
The broker is exactly the additional moving part I was referring to.
Pubsub over a broker is more complicated than RPC operationally, whether you're already using the broker software elsewhere or not.
Especially when you're looking at the RPC server living inside an erlang VM that's already really good at handling things like load shedding of direct connections.
To be clear, this article is talking about implementing your own synchronous-RPC-request server, i.e. a network service that other services talk to through an API over some network wire protocol, to make requests over a socket and then wait for responses to those requests over that same socket. This article assumes that you already know that that's what you need. This article then offers an additional alternative to the traditional wire protocols one might expose to clients in a synchronous-RPC-request server (RESTful HTTP, gRPC, JSON-RPC over HTTP, JSON-RPC over TCP, etc.); namely, mimicking the wire protocol Redis uses, but exposing your own custom Redis commands. This choice allows you to use existing Redis client libraries as your RPC clients, just as writing a RESTful HTTP server allows you to use existing HTTP client libraries as your RPC clients.
The alternative to doing so, if you want to call it that, would be to write these custom commands as a Redis module in C. But then you have to structure your code to live inside a Redis server, when that might not be at-all what you want, especially if your code already lives inside some other kind of framework, or is written in a managed-runtime language unsuited to plugging into a C server.
Or think of it like this: this article is about taking an existing daemon, written in some arbitrary language (in this case Elixir), where that daemon already speaks some other, slower RPC protocol (e.g. REST over HTTP); and adding an additional Redis-protocol RPC listener to that daemon, so that you can use a Redis client as a drop-in replacement for an HTTP client for doing RPC against the daemon, thus (presumably) lowering per-request protocol overhead for clients that need to pump through a lot of RPC requests.
I do realize that you're suggesting that you could use Redis as an event-bus between two processes that each connect to it via the Redis protocol; and then use a "fire an async request as an event over the bus, and then await a response event containing the request's ref to show up on the bus" RPC strategy, ala Erlang's own gen_server:call messaging strategy. All I can say is that, due to there being three processes and two separate RPC sessions involved, with their own wire-protocol encoding/decoding phases, that's likely higher-overhead than even a direct RESTful-HTTP RPC session between the client and the relevant daemon; let alone a direct Redis RPC session between the client and the daemon.
Event bus is the name for this network architecture. And if you're trying to replicate what synchronous client-server RPC does in a distributed M:N system, it's what you'd have to use. You can't use at-most-once/unreliable PUBSUB to replicate how synchronous client-server RPC works, as a client might sit around forever waiting for a response that got silently dropped due to the broker or a backend crashing, without knowing it. All the queues and ACKs are there to replicate what clients get for free from having a direct TCP connection to the server.
(Yes, Erlang uses timeouts on gen_server:call to build up distributed-systems abstractions on top of an unreliable message carrier. But everything else in an Erlang system has to be explicitly engineered around having timeouts on one end and idempotent handling of potentially-spurious "leftover" requests on the other. Clients that were originally doing synchronous RPC, where you don't know exactly how they were relying on that synchronous RPC, can switch to a Redis-streams event-bus based messaging protocol as a drop-in replacement for their synchronous client-server RPC, because reliable at-least-once async delivery can embed the semantics of synchronous RPC; but they can't switch to unreliable async pubsub as a drop-in replacement for their synchronous client-server RPC. Doing the latter would require investigation and potentially re-engineering, on both sides. If you don't control one end — e.g. if the clients are third-party mobile apps — then that re-engineering might even be impossible.)
This is how I read the article. It was about how to implement network protocols in elixir and here are two of them: redis and msgpack.
Having an elixir based redis server is not the same piece of the puzzle as having elixir simply talk to a redis server. For one, the elixir based redis server can have arbitrary rules around keys and values that are not supported by redis. (Same said for a redis server written in C or python or rust or...)
This approach lets you store all keys and values in a dict, b-tree, sqlite or postgres etc. Want to store each value in a flat file? Sure, now you can. Only you know if this is actually useful.
At least, this is how I made sense of this.
I think the article was instead just using the term "Redis server" to mean "any server that speaks the server side of the client-server protocol that redis-server speaks" (without implication that it stores data ala redis-server) — in the same way that an "HTTP server" is "any server that speaks the server side of the client-server protocol that HTTP servers speak" (without implication that it serves HTML files, generates directory indices, and supports per-user multitenant shares, ala default-configuration Apache.)
Note that the command set of Redis isn't part of the Redis protocol. A "Redis protocol" server could have an entirely novel set of commands, none of which have anything to do with keys or data-structures. It's just another way of exposing an API to clients. Your application-layer protocol over the Redis wire protocol could be "an API for triggering webhooks", or "a group-chat software protocol ala IRC/XMPP", etc.; and in none of those cases do you need to implement GET/SET/DEL/etc., or to describe your own use-case in terms of GET/SET/DEL/etc.†
The only thing the Redis wire-protocol necessitates, IIRC, is that each command start with a verb; that verbs consist of ASCII characters, with a certain maximum length; and that each verb have a fixed "schema" for the members of its parameter list, that pre-determines the encoding a client should use to send an instance of that command over the text or binary wire-protocols, without any connection-time schema discovery.
And yes, most Redis clients do have some way of sending custom commands to the server, with the schema for those commands specified at runtime (at least over the Redis text protocol), even if they don't have syntax sugar for doing it the way they do for the redis-server built-in commands. Even if a Redis client's aim is only to talk to redis-server deployments, they still have to support custom commands, because individual redis-server deployments are extensible with https://redis.io/modules that expose arbitrary commands, and client libraries can't possibly know about those modules at compile-time. So they have to support potentially any command at runtime, somehow or another.
-----
† Mind you, just like HTTP has "REST" (which basically means "using the default HTTP verbs for analogous purposes in your own API, instead of totally abusing theirs semantics or inventing your own verbs"), there could be a similar convention on top of the Redis wire protocol, where you implement your API in terms of the built-in redis-server verbs GET/SET/DEL/etc. Then you could use the full syntax-sugared default commands built into Redis-protocol clients, to talk to your server, instead of needing to rely on the runtime custom-command support. However, unlike with REST, I don't think this use-case is very useful — the schema of redis-server's built-in command verbs has pretty tight tolerances, and doesn't allow for too many use-cases that aren't just "building a data-structure server."
Thanks for this beautiful analogy. I'm familiar with redis and elixir and found the article interesting but I didn't quite understand the "why" and now this makes sense. Highlighting it here for others as well.
You should write technical book or something!!
For example, to send "SET my-key hello" to Redis, this is a 3-element array (the initial *3) with elements of length 3, 6, and 5. So you'd send:
*3\r\n
$3\r\n
SET\r\n
$6\r\n
my-key\r\n
$5\r\n
hello\r\n
(^ all of those would be joined together, this is presented on multiple lines for readability)Decoding lengths from their textual representation is inefficient in both space and time, and the various \r\n add to the request and response sizes without adding anything useful to a program that uses it (it is of course useful for humans typing those commands, but who does that?).
Even using something as common as protobuf would likely be more efficient, with smaller requests and responses – although object creation can lead to memory pressure for complex messages.
But really, since Elixir leverages Erlang, why not have a basic common protocol with pattern matching on binaries? This would make a lot more sense than using a text-based protocol, in my opinion.
Redis took a middle ground and [ascii]-length prefixed responses (though I don't think it was initially like this). That's a reasonable and likely necessary compromise. A bit similar to how HTTP can include a content-length. You'd hate to be scanning large payloads for that trailing \r\n.
So Redis terminates with \r\n (which isn't strictly necessary for all message types) and ascii encodes the types and length prefix. Everything else is just bytes.
Here's (1) a paper that looked at text vs binary protocols. They looked at it in the context of MonetDB though, which has, by far, the worse text-based protocol I've ever seen (so much so that even their official drivers have (had?) plenty of protocol-related bugs)
(1) - https://15721.courses.cs.cmu.edu/spring2018/papers/14-networ...
It's true that the parsing time is small compared to the network round-trip, but Redis being single-threaded means that while it's parsing a large nested command and allocating memory as it goes through it… well nothing else is being processed.
I'm pretty familiar with the costs of parsing this protocol, being the original author of the phpredis extension. Written in C, it's a high-performance client whereas pure-PHP clients can't use the same zero-copy tricks or careful memory management in general to avoid repeated calls to strlen/malloc/memcpy/free (in that order). In Webdis – another high-performance Redis tool I wrote – protocol parsing with Hiredis is similar to HTTP parsing in the sense that it's where most CPU cycles are spent.
I'm not really sure why Salvatore chose to keep the text-based protocol when his focus was so often on performance. I'm not aware of any attempts made by him or other people to change the protocol, but it would be interesting to see how much of a difference this could make.
An average latency of 17ms to handle a simple request with the resp protocol is a red flag it self that something is wrong. Seconds to have an http response returned is a problem. All this article is showing is elixir and the libraries used are incredibly slow. Calling HTTP slow, based off the tests done in this article is just wrong, because the whole test is slow and not testing the right thing.
Micro bench marking of the parsing functions would be the right thing to do here.
really someone should profile the server and see what is taking seconds to respond. it wont be the parsing related functions
parsing http really isnt that big of a deal. its a simple protocol, the action on line one, headers split by : and the body. resp has a simmilar amount of work too, uses multiple lines
The redis protocol and Msppack are both only a a hundred lines or so for a parser. Meaning you can build your own from scratch in a new language if one isn’t supported.
Its also stupid fast.
Compared to protocol buffers which can be extremely complicated to grok on the binary level.
I built a more robust API RPC for Python here based on Redis and MsgPack: https://github.com/hansonkd/tino
And on one hand, sure asn.1 exists, on the other there's been some issues in for example nfs. I don't think you want asn.1 today - but maybe cap'n'proto.
But i think redis is probably an interesting approach.
Not sure how many of these beyond whatever/http2 have a sane pipelining+authenticated encryption story? I suppose you could "dictate" security at the ip level via vpn/wireguard or something. Or use Unix sockets.
I also think there were some nfs security issues relating to asn.1
https://www.infosecmatter.com/metasploit-module-library/?mm=...
I guess that (and related errors in openssl and bouncy castle) might not be due to problems with asn.1 per se, but rather difficulty in writing safe libraries for parsing in C.
(disclaimer: I'm one of the NATS maintainers :) )
That seems like a weird claim. Protobuf on the wire only has four types: short fixed-length values, long fixed-length values, variable-length values with continuation-bit encoding, and length-prefixed values, where the length is continuation-bit encoded.
Msgpack has 37 different wire types!
I’m showing that it is just as easy to use a different protocol than HTTP.
Instead of reaching for an HTTP framework to do RPC, I’m showing that Redis can be just as easy.
And because there are decades of best-practices and standard ways to do very common things like authentication, caching, signaling errors, load-balancing, logging. Maybe a homemade RPC doesn't need any of that, but I would bet that at some point it will need at least some of it. And now you have to implement it all yourself.
This is cool though and kinda fun to think about!
Essentially either you want RPC, or you might want caching. If you want to put a varnish cache in there, probably stick with http. If you know you're doing (very high thro / low latency) RPC - this might make a lot of sense.
So it's not so much if you want to memorize (cache), but if you can* (and/or should).
I'm afraid I'm repeating myself - but I really whish people would read all of Fielding's thesis - it's got all kinds of reasonable architectures in it - even if it argues hypermedia/hypertext applications are well served by REST.
These benchmarks, while dated and haven't been tested with the newly released updates to address these dynamics - might interest some folks.
In this case, both HTTP and Redis use TCP. If HTTP header and data fit in 1 MTU, there should be no difference. No way searching for 2 CRLF that too in 1 MTU data size be 100x slower
But you won't miss all the http reverse proxy ecosystem.
https://performance-dot-grpc-testing.appspot.com/explore?das...