Clojure's edge on Node.js
dosync.posterous.com
dosync.posterous.com
So node would probably outperform the given clojure example by 13.5K req/s vs. 8.5k req/s (YMMV).
It is also worth noting that clojure library used (aleph), does not seem to support streaming the response.
The most interesting aspect is clojure's support for futures, and Javascripts lack of such. If you're willing to trade speed for cheap syntactical pleasure, clojure has an advantage here : ).
edit: Who down voted this / why?
Also I note that your code is very verbose for managing multiple cores while the Clojure version is pretty much identical to the original Node version.
So much for cheap syntactical pleasures, eh?
Anyway, I think you are mistaken if you believe a clojure http server could outperform it's node counterpart. I mean node is really just a set of air-thin Javascript bindings to some pretty good C code / native system calls. Node's http parser is also written in C.
Shows node to outperform nginx for serving a static 1 MB buffer.
Anyway, you are correct - a real benchmark between node / clojure would be needed to proof my argument. But I think reasoning about the underlaying machinery is a better use of my time in this case. Bugs aside, there is no reason to believe Clojure could perform faster if it ends up doing more work under the hood. And from what I can tell, there are many more layers of abstraction in between clojure's http implementation and the native system calls, than there are for node.
Since V8 does not have native continuation support, realistically it's going to devolve to libraries either doing CPS transformation (http://gfxmonk.net/2010/07/04/defer-taming-asynchronous-java...) or monads (http://matthew.yumptious.com/2009/04/javascript/dojo-deferre...), which both involve performance penalties.
A fair comparison on say a 4-core machine would be 4 JVMs running Clojure webservers with some front-end, and 4 node.js instances running one of the library front-ends, with the benchmarked page written with one of the CPS-transforming toolkits (coffeescript, narrative JS, jwacs, etc.) or maybe monad toolkits if/when somebody decides to write one.
Check out http://github.com/kriszyp/multi-node or http://github.com/extjs/Connect for examples of web servers that will blow Clojure out of the water in this benchmark.
When leveraging multiple cores, NodeJS is closer to nginx in terms of performance than it is to Jetty or Django. (Of course, nginx uses almost no memory, and NodeJS has this big honkin VM juggling JavaScript contexts, so it's not as space-efficient.)
Also, don't forget that GC penalties are going to be paid by all threads in a process.
As to GC, that presumes that you are using GC. For a request/response style server application, there are more optimal techniques than GC: make sure all request-associated allocations come out of a per-request heap, and you can free that heap en masse, in one go, once that request is done with.
For everything else ... i.e. most startups I see on HN ... you really have to take the CPU-bound or synchronous logic out of Node.js, otherwise it sucks big donkey balls.
People striving for the ultimate development experience have really forgotten what it's like to have brute processing power. Web apps like PlentyOfFish which had 1.2 billion page views/month in 2009 ran only on 2 balanced web servers ... and this considering that on a dating site everything's dynamic, you really can't have too much caching.
Clojure might turn out to be the best of both worlds ... a productive programming language coupled with the brute CPU-bound performance that the JVM can yield.
I'm not completely sure about this, but I remember hearing that SpiderMonkey is more suited to using threads, and that's why CouchDB uses it rather than V8.
So I'd regard both as "things to watch" or bootstrapping solutions, not something I'd use for my Global Thermonuclear War management system.
If you're interested in playing around with tech and having fun, then these technologies are great. But for business I would avoid them for now.
I'm hoping they gain some traction though.
I'm certainly not claiming it's flawless, or that there won't be any breaking changes in the future. However, I think this is a great example of one of Clojure's strengths: even brand new projects tend to have solid, well-tested foundations.
At the very least some hackers out there are able to make it scale quite well. Does anyone have a comparable anecdote for Node.js (honest question)?
http://amix.dk/blog/post/19490#Plurk-Instant-conversations-u...
In Clojure we don't need to use callbacks. This means for common things like talking to databases, we don't need them to have asynchronous interfaces. That's because we have really fantastic primitives in the language itself for dealing with concurrency. This code runs twice as fast as the Node.js counterpart - probably due to the excellent perf of Clojure coupled with leveraging multiple cores.
"Blocking (or possibly blocking) system calls are executed in the thread pool. Signal handlers and thread pool callbacks are marshaled back into the main thread via a pipe."
What you ought to have is multiple async accept calls outstanding, and as they complete they are entered into work queues. Worker threads (i.e. the thread pool) pulling work of those queues should steal work from other queues when their own queue is empty.
See e.g. http://www.bluebytesoftware.com/blog/2008/09/17/BuildingACus...
Pulling items off a queue, either by its associated worker thread or another worker thread stealing its work should be lock-free.
As I mentioned elsewhere, GC is not necessarily (or even often) the most efficient approach for request / response style servers. A more optimal approach is a heap associated with the request which can be freed in a single go when the request is done with, with all allocations associated with that request (i.e. that don't need to persist between requests) coming out of that heap.
You can even design a GC around this principle: have one GC heap per worker thread, and collect it after every request has been processed. There should be little or no roots for this heap associated with the worker itself, which should be (very) low down in its call stack after it is done with the request. If you have write barriers for any mutations to inter-request (shared) state, you can trace those to find out which bits of the worker thread's GC heap you need to keep (copy out). Then you can simply zero the GC heap and reset the free pointer. You can make your write barriers smart so that they are associated with that worker's heap, so you don't have to wander all over the shared heap looking for roots.
Are you talking about thread-safe queues implemented using atomic instructions? Those aren't free - how do you think they're implemented at the hardware level? The main advantage of atomic instructions over locks is removing the possibility of waiting on a preempted thread holding the lock (lock-freedom). They also have lower overhead than making system calls to provide locking. But the equivalent mutual exclusion logic (and contention penalties) just get moved down to the chipset level - now instead of waiting on other threads, you're waiting on other cores/CPUs.
"What you ought to have is multiple async accept calls outstanding, and as they complete they are entered into work queues. Worker threads (i.e. the thread pool) pulling work of those queues should steal work from other queues when their own queue is empty. See e.g. http://www.bluebytesoftware.com/blog/2008/09/17/BuildingACus...
That article is horrible. Please do not follow the author's advice.
Besides the fact that the code deadlocks (see the first comment on the article), it's also easy to see that if one thread starts generating all the work the "solution" is going to become a single global queue (the "local work stealing queue" of that thread), with all the other threads looping through the global queue and then all each other's queues just to reach there!
The one good thing about that article is that the throughput gains on the toy benchmark (which doesn't deadlock or spawn tasks non-uniformly) nicely illustrate my point about the expense of contention even if using atomic instructions.
What the code in the article attempts to do is alleviate contention by partitioning the tasks among several queues. The problem is that if the work is not distributed uniformly among the queues, some threads will be left idle. The way to overcome that is to fake a global queue by having some way to synchronize the partitioned queues. The optimal solution to how to do this depends not only on the particular system you're running on, but also the pattern of work spawning by the application. And all this depends on being deadlock-free (something not managed by the article)!
Why would you ever go through something so horrible for an HTTP server? Multiple node.js processes are much simpler and more efficient.
PS - The "GC" you're describing is called CGI.
"The problem is that if the work is not distributed uniformly among the queues, some threads will be left idle" - this is why you use work-stealing queues! The very nature of work stealing queues is that the worker threads aren't left idle - they steal work from other threads' queues.
And CGI is not anything like the GC I talked about - if you have a process per request, where are you going to put your shared state?
But much of this discussion is besides the point. Don't forget, the OS scheduler is at its heart an event dispatcher when there are more runnable threads than CPU cores. The thread stack is little different than context provided to a triggered event. You want to have the same number of runnable threads as CPU cores in order to avoid the kernel cost of a context switch. You can do that by having multiple single-threaded processes, or multiple threads in a single process. While neither choice of partitioning affects the degree to which you can use an eventing style to serve requests, one - the separate process model - makes it much harder to share state. And therein lies the reason why I believe that optimal performance lies in threads, rather than processes. There are other good reasons for using processes instead - but it will be at some cost to efficiency.
function do_something():
socket = wait_for_a_socket()
data = read_everything_from_socket(socket)
process_data_long_and_hard_with_lots_of_io(data)
write_to_disk(data)
return_result(socket, data)
Where, as my function names imply, further code may be called that does things like talk to databases, or wait for other data, or any amount of other I/O, and you may not have to write a single asynchronous callback. Why? Because unwrapping code into continuations and managing them in the runtime is a trivial compiler transformation when you design your language to work like that from day one. Like Erlang, or, in this case, Clojure.The fact that Node.js requires you to manually shatter your code into teensy-weensy little fragments and manually wire it back together is a hack which should not be mistaken for a feature. Javascript requires you to do that, because it's an Algol language and just doesn't work any other way. If that's what you love doing, great, but that's an awful lot of time and effort spent on writing plumbing (and debugging plumbing, and debugging nontrivial asynchronous plumbing in a mutable language isn't in the worst tier of programming tasks in the world, but it's solidly in the second...) that you could have been spending on writing code that actually solves customer problems.
(I am aware of the libraries that pretend to help this. They are a joke compared to working in a language that actually supports this. I can call functions and send messages and read files and read from databases and write functions to abstract all this and I don't spend one second wondering how I'm going to wire all the pieces together at runtime. No amount of code slathered over Javascript can match that, short of an entirely new language that compiles into Javascript. (Which is inevitable. And it will be hailed as a brilliant breakthrough.) All of the libraries I was pointed to last time I brought this up use the exact same obvious hack, which helps with the case of stringing a handful of teensy-weensy fragments of code onto a single string but can't handle anything more.)
I set up a Rack adapter through Jetty 7 on JRuby that works exactly like this. The SelectChannelConnector is event-driven I/O, and the QueuedThreadPool handles requests in parallel.
The result? 6k req/sec on my laptop for a simple Rack app, basically the same as what's shown in these examples, since my machine is quite a bit slower, and I've clocked 10k/sec on faster machines.
(By the way, if you're interested, the Rack handler is in a gem called 'mizuno', and is at http://github.com/matadon/mizuno)
Pro-tip: This time, you were looking for the word "then". But "than" really exists too, and is even used sometimes (when comparing things)!
Yeah, downvote away.
But you're correct. :)
And I also don't downvote anything except outright spam; everybody has the right to an opinion, whether or not I agree with it.
[1] http://www.ecma-international.org/publications/files/ECMA-ST...
[2] http://www.whatwg.org/specs/web-workers/current-work/
[4] http://odetocode.com/blogs/scott/archive/2007/01/08/javascri...
Don't forget to add a non-Javascript version for people without Javascript (rare) and bots (common).
Actually, using something like nginx-pr (a-la apache's apr) would be much better solution than this bunch of hacks.
But people never listen. ^_^
I don't know how clojure works.
I know how clojure works.
...
Code is data :)
I got to the point of understanding and enjoying clojure's syntax very quickly.