A lot of websockets in Haskell
blog.wearewizards.io
blog.wearewizards.io
Haskell will use less memory on this task. Since it has a shared heap, it doesn't have to allocate a small heap per process, and thus it is expected that it has lower overhead. Furthermore, static typing means Haskell needs less type tags and this tend to make it win. As the system grows in complexity, these things tend to even out a bit more but I would still expect Haskell to use about half the memory of Erlang.
The Elixir numbers for Phoenix sounds off. At 83765 megabytes and 1999984 connections, that is 42 kilobytes per connection. That count is about an order of magnitude over what I would expect it to be. How much of that memory is kernel allocated network buffer space, and how much is buffer space in the Erlang runtime? A "raw" process in Erlang is around 1.5 kilobytes nowadays, including stack and heap, so where do the additional 40 get allocated to? I don't think we have 20 extra processes per connection for some reason :) Definitely something to look into.
Tsung is an old application. It isn't really written in ways that makes it efficient at the network level and this shows its head in this benchmark. Furthermore, Tsung does more work than the broadcast in Haskell, so it is expected the load generator will give up long before the server. Again, measure the amount of memory allocated by the kernel and by the userland process in order to determine if it is one or the other you hit first. Still, I would expect Tsung to be the culprit.
In addition, the classic PubSub pattern here screams the Disruptor pattern (which I first heard about from Trisha Gee and Martin Thompson). The Erlang runtime has no direct good support for this kind of pattern, so you have to opt for things such a ETS to simulate it. It'll work albeit at an overhead.
For web server benchmarking, only wrk2 by Gil Tene does things correctly. Everything else usually does coordinated omission:
Imagine you have 10.000 connections. Each connection is doing 3 req/s. Let's say one connection blocks for 1 second, which means that 2 req's should have fired on that connection "in between". wrk2 will count those two as being "late" whereas most other load generators won't count at all. This means a framework can opt to "stall" some connections in order to get better performance and fewer bad results in the upper latencies.
As an example, here are the Erlang/Cowboy numbers for such a test in wrk2:
Latency Distribution (HdrHistogram - Recorded Latency)
50.000% 6.33ms
75.000% 10.23ms
90.000% 13.73ms
99.000% 22.37ms
99.900% 31.26ms
99.990% 38.62ms
99.999% 45.09ms
100.000% 49.60ms
Whereas the Haskell wai framework returns Latency Distribution (HdrHistogram - Recorded Latency)
50.000% 1.71ms
75.000% 2.64ms
90.000% 14.92ms
99.000% 38.75ms
99.900% 742.40ms
99.990% 985.60ms
99.999% 1.03s
100.000% 1.05s
Note how the median latency and the 75th percentile is better for Haskell, but that it occasionally stalls requests for quite some time, probably due to a GC pause or some other cleanup that happens and then in an unfortunate moment all has to happen at the same point in time.If you go look at typical benchmarks their latency reporting is way off compared to this, which is a surefire way of knowing they did not account for coordinated omission.
Mind, when benchmarks disagree, the trick is to explain why this happens. It often leads to an insight in design difference.
[1] https://github.com/TechEmpower/FrameworkBenchmarks/issues/12...
I'm also curious if dependently typed languages like Idris, which presumably must be able to have runtime access to type information, handle this stuff.
For functions, because a function that takes two arguments and returns a value (a -> a -> a) has the same type as a function that takes one argument and returns a function that takes another argument that returns a value (a -> a -> a), the arity of functions is stored in the tag.
Some of these tags are eliminated by inlining but if you sit down and read some typical Haskell output you'll see a _whole lot_ of tag checks.
Source: spent a lot of time reading GHC output and writing high-performance Haskell code.
I wrote a mail server in Haskell that, because of a bug was leaking connections. Didn't use nix, so I had the external forkIO explicit, and used a custom sockets interface, that creates a buffer for every connection. I optimized almost nothing, this was an earlier version of it.
Anyway, every time the server would reliably get to about 700k open connections before getting killed at my linode machine.
500k users per machine is great if they are mostly idle. This is the use case of WhatsApp, and their stats are[1]:
> Peaked at 2.8M connections per server
> 571k packets/sec
> >200k dist msgs/sec
Not every app is meant to have mostly idle users. Can a real-time MMO FPS be done in the Haskell server is the question (lets limit each user's neighbourhood to the 10 closest players).
I'd be very interested in the other corners of the envelope for the Haskell server: A requests/second benchmark over websockets, with the associated latency HdrHistogram, like [2]
[1] http://highscalability.com/blog/2014/2/26/the-whatsapp-archi...
[2] http://www.ostinelli.net/a-comparison-between-misultin-mochi...
Haskell code with tester is here: https://github.com/bbangert/ssl-ram-testing/tree/master/Hask...
(In some cases it would take some client churn to get the leak to occur, but clients come/go, a server has to survive this basic fact)
One diagnosis was that it was a memory fragmentation leak, since sure enough when I ran some debugging it reported fragmentation losses at the end. Some tweaks to the initial stack size and how much stack each increment should be done remedied it for the most part, but resulted in it taking quite a bit more memory (I believe I had to set the initial stack to at least 4k).
Using -N4 or -N6 (to turn on multi-core/cpu use in haskell) increased memory consumption quite a bit, as Haskell will now start juggling all those threads across real OS threads with its M:N scheduler (Just like Go has to do with its goroutines). It was drastically more memory efficient with -N1, even though that loses the multi-core utilization.
So far the most efficient implementation I've tested is of course... in C. Where a simple websocket echo used 9.5kb of memory per connection (including SSL, and with TCP kernel buffers set to 4k send/recv). I'm not really eager to write C code though, so we're using Python+twisted where a plain WS connection will take about 16kb per conn, or 23kb with SSL.
In this article's test, the kernel buffer for TCP send/recv was set to 1k, which I guess is fine if your payload will frequently be under that. But if you regularly are sending payloads exceeding that, the amount of wakes to keep shoveling data into the kernel buffer is going to be rather expensive.
As the Phoenix people note, its useful to keep in mind what you need to do, rather than merely how many connections you want to hold open.
Perhaps a microservces implementation with each step (handshake, broadcasting, http overhead) having their own cluster and communicating through channels.
To anyone else who uses nixops, why does nixops use this state file? Can't they use labels like Ansible to identify deployed machines? Can't you just require the user to provide their own secrets instead of auto-generating a new keypair for each machine?
Thanks for the writeup!
Really cool example! Just FYI, pre-emptible VMs would only be $0.06 (vs. the bid of $0.10 for AWS Spot Market). I mention because the author talks about cost being a concern.