C10K 2012: Erlang Wins, Go a close second. Java, Haskell, Node & Python fail
github.com
github.com
I see the Tornado script forks processes, for example, but it doesn't experiment with different values (or PyPy?). The Java code appears to be run with default VM settings, which pretty much always deserves some tweaks to run in a server environment.
Those are just the two I'm most familiar with. The rest seem to have similar faults. Plus, it would be a more interesting test if the server had to perform some kind of work. Simply echoing the request is not a typical usage and could really bias the results in favor of setups that would fail on a real project.
I better title might be, "Erlang wins unrealistic test over other VMs in their default configuration."
I have benchmarked the ws4py code in the test on that machine, and it was much slower than the tornado code - quite obvious since it ran in a single process thus on a single CPU.
BTW The default number of processes when you fork is the number of CPUs, which sounded like a fair estimate to me, not knowing the target machine, so I left it.
This is always the response to any benchmark where one's pet technologies don't win. If you have a magic quadrant of tuning, put it forth. If not, you have said nothing that counters the results.
Maybe it doesn't call it out explicitly - but in my view that is a reasonable enough test.
I'm pretty sure all the platforms could support an FFI binding or equivalent to an optimized epoll C implementation - or hours of tuning.
This one in particular is clearly incomplete, especially when things like Haskell, Java, and Python are represented by a single framework, and not even the best or most optimized ones. For example, why not use Yesod/Warp, BlueEyes, and Tornado on PyPy, respectively, for those instead?
Also, the submission title is editorializing flame bait. The study's author didn't make any claims whatsoever about 'Haskell', 'Java', and 'Python', etc. but rather about Snap, Webbit, and ws4py.
It's always possible for the critics to go out there and do something better, rather than limiting themselves to pointing out flaws.
Maybe in an ideal world, quality benchmarks would be commonplace. But in this world, doing good benchmarks in response to everyone's bad benchmarks is a heavy burden.
So maybe the study's author only cared about webbit, ws4py, and Snap, and that was the implicit constraint of his test. If so, that's fine, and the submitter mangled the author's intent with an overly-editorialized, senationalized title.
But if the author was actually using webbit, ws4py, and Snap as proxies for all of Java, Python, and Haskell, then no one should be surprised that it gets quickly shot down. There's value in critics debunking BS quickly, even if they don't provide a better alternative. Absence of bad knowledge is better than presence of it.
Nobody has to accept this benchmark as indicative of anything if they don't find it robust enough.
I'm really impressed that erlang didn't drop a single connection. I suspect if you ran the same test with just erlang 5 times, you'd see it start to behave like go.
I don't think it would. Erlang was designed for things like this exactly and has many many many years up on the competition (Go in this case). I'm sure Go will get there, but by then, with Erlang's further threading improvements coming in R16 it will also get better.
I really like Go and can believe that Erlang performs well. For the java example.. wtf is "webbit"? I've programmed Java for 10 years and never heard of that. Why doesn't he just find a way to run on Jetty or Tomcat like everyone else?
I'd believe for an example like this that native languages can outperform java, because it's basically a ton of no-op web requests. But Java underperforming Node and frickin Python? Python's a great language but famously slow. No way is this study legit.
It's a benchmark of specific websocket implementations, run with default configurations, which happen to be in different languages.
Drawing any conclusions about the merits of each language based on the results would be very foolish.
Basically, benchmarking is hard, and if you threw together a quick benchmark over a weekend, there is a good chance it might not be representative.
Erlang and Go doing great is no surprise. Java didn't do too well, but the implementation didn't really use the most performant tools available.
The real loser here is Node.js. Nearly as long as the Java example, but the worst performer in the real metrics. So much for "making concurrency easy."
Erlang is automatically threaded by nature so it has some inherit scaling built-in (so the code benefits from threading as long as it runs correctly).
The Go code is set up to spin up threads inside of ListenAndServe thus gaining the benefit of splitting up IO.
The Haskell code has a thread specifically for garbage collection, thus utilizing the second core.
The Node.js code could be using cluster (or threads/fibers more directly) but isn't for some reason. It also seems this was using the websocket npm (some unknown code running on the stack!). For a valid test the websocket code should be written in JavaScript directly in the test itself.
Edit: Researched and Go actually uses ListenAndServe which creates threads on-the-fly but the m1.medium is bound to 1 CPU (it still benefits from threading due to IO).
> The Go code is set to spin up threads to accomodate the number of CPUs available on the system (2 logical cores on m1.medium).
m1.medium is a single virtual cpu (it is a vm, so no 'cores' to speak of). c1.medium has 2.He doesn't mention increasing somaxconn, which means it probably has a default of 128 (I don't think increasing tcp_max_syn_backlog is needed since he's using syncookies). When creating a new connection every 1ms it only takes 128ms for the backlog to fill up, which is not that improbable with GC and JIT pauses. Java should be faster than erlang most of the time, but the pauses can kill you if you don't handle them gracefully.
If the benchmark had been about "normal" http requests, I would suggest putting HAProxy in front with a reasonable maxconn - then HAProxy will hold on to the connections until the application starts accepting again.
To say it bluntly, "put your money where your mouth is".
Throw uWSGI behind Nginx (with it's new websocket support), tune it a bit, and I wouldn't be surprised to see it "pass" and perhaps even be competitive.
Read the timing numbers with a reasonable portion of salt. It's EC2 we're talking about here.
(I know it's then close to 'dedicated' hosting, but people running benchmarks also want the quick setup and discard of virtualized instances. This would be a hybrid offering that ensures their cloud services is always chosen for such benchmarking comparisons. Of course this offering would not be good for cross-cloud comparisons, because it's not representative of their usual offerings.)
That's called EC2 cluster compute.
- Using VPC you can specify that your instance should run on dedicated hardware [1]
- Cluster Compute instances [2] most likely (see [3]) run on dedicated hardware too
[1] http://aws.amazon.com/dedicated-instances/
[2] http://aws.amazon.com/ec2/#instance
[3] https://forums.aws.amazon.com/message.jspa?messageID=238197#...
If you were to conduct a test like this, where would you go for more reliable performance from hardware?
any framework, any language, hell any hardware is either generalized to deal with many problems or specifically designed to battle one.
you sure can write fast C program that does exactly this benchmark well - what will that prove? nothing at all..
why is it so hard to accept the fact that some tools can provide good enough results without too much tweaking while other tools may provide better result with more time spent achieving it?
For example, looking at the raw data for Haskell and java, and adding up connections, disconnects, crashes, and timeouts doesn't give anywhere near 10k. What happened to the rest of the connections? Were they simply not attempted? Was the port closed? Or am i just missing a relevant field in the output? It's not clear to me from the description.