RethinkDB internals: making coroutines fast
rethinkdb.com
rethinkdb.com
As you mention malloc()ing stacks: you might want to allocate them using mmap() and MAP_ANONYMOUS instead. You can map the adjacent memory pages with appropriate protection to prevent stack overflows and as a result, memory corruption. I'm not aware of any drawbacks (malloc itself typically uses mmap above a certain size) but it certainly beats hoping your stacks are big enough. Less of an issue on 64-bit environments of course.
I understand that it is difficult to change JavaScript, yet, when you create a whole new framework, say, nodejs, why not providing coroutines?
Does it need a language construct to be efficient? Maybe, then have a look the Icon programming language, where coroutines rule,
Lua has one of the best coroutine implementations I've seen. Scheme has continuations which are convertible to coroutines, while Lua has coroutines which are convertible to (one-shot) continuations, with community insight on their use.
Some e.g., Asana have added fibers to Javascript (in their case V8): http://asana.com/blog/?p=49
The previous post from RethinkDB is quite interesting in terms of motivation for coroutines: to me, this superficially seems like SEDA. Here, instead of stages in the pipeline (thread pool per stage, first stage being processing events from epoll/kqueue, thread per core in each pool, each thread holding state machines, communication between threads via a queue) you are using coroutines for clearer code.
I agree about your point regarding IPC, but that's something they can always add later. It's unfortunate that they've gone the callback route instead of innovating like Asana has done with coroutines.
Add some form of IPC as an afterthought? Sure. Create powerful primitives for IPC as done in Erlang/OTP? No.
It's not a bad thing that data has to be copied to be shared between Erlang-style processes. That is part of the reason why things "just work" when you move that process to another machine, or data centre. If the implementation can be smart and "cheat" by sharing data in the same process that's fine, but it is an implementation detail and not a property of the system.
The way Erlang works is that you have a single event loop and threads are spawned to handle i/o and such (the n:m threading model, N green threads are mapped onto M OS threads). When a process blocks its execution is suspended.
Node employs an n:m threading model as well except "processes" in node are just functions. One big difference is that if you make a blocking call in node the whole process is blocked. There's only an event loop, no scheduler or anything that OS-like involved. The Erlang model is clearly superior imo, but Node is far more accessible (for better and worse).
If you don't care about performance, sure (there's nothing wrong with not caring about performance e.g.,: low volume web applications, where server side Javascript can shine, IMO). However, I can give you a command line you can run which will plot a very nice graph of throughput vs. number of selector threads in a specific system that I've been working on.
Yes, it will bring you back to "concurrent madness". What you _want_ is having primitives that make it possible to deal with this madness, not handwave it away. In Erlang they're actors themselves, optimized for efficient delivery of message within a single node and remotely, supervision of processes that fail, ETS, mnesia; in Java they're Doug Lea's beautiful concurrent collections (java.util.concurrent). You're assuming multi threading pthreads or synchronized/notify/wait (Java before java.util.concurrent).
Note that there are two models for this: one is Erlang's as well as traditional UNIX IPC -- message passing; the other is shared memory multi-threading with concurrent and lock free collections (Java), which also goes nicely with the idea of minimizing mutable state (Haskell, Clojure, Scala). One is good for one kinds of applications which optimize for worst case latency (Erlang shines there), the other is good for another, which optimize for average throughput (JVM shines there).
http://en.wikipedia.org/wiki/Grand_Central_Dispatch#Examples
I don't know a lot about them though, and having an extra benchmark against a similar SSD installation of MySQL or Postgres with the same select query that achieved 1.5 million QPS would be pretty interesting.