Why thread-based application parallelism is trumped in the multicore era
ibm.com
ibm.com
Each process has its own memory space and its own little bit of the work load to complete. One process can crash or throw and exception and the others keep on going. Having no shared streams or shared data containers to worry with (mutexes, locks, etc) is just wonderful.
We call it poor man's parallelism and some guys who have done a lot of threading make light of it. It's so simple (compared to threads) that it seems like a naive approach. But it performs so well that it's hard to argue with the results.
We've found that this will beat a 'better' parallelism even if we need to do a lot of extra or redundant work due to the lack of IPC or reduction due to the no-overhead approach, although this will of course be highly dependent on the problem at hand.
I think you're dismissing the state of the art on multiple fronts, as is the article. From single-variable STM (of which Haskell and Clojure both have excellent implementations) to battle-tested and well-understood concurrency primitives in the Java standard concurrency library, multi-threaded programming is more approachable and performant now than it has ever been.
But the article is oddest in that it seems to hold up Erlang as a way forward, but Erlang is just a different model built on top of thread-based concurrency. If the argument is the old pthread-ish model of "1 thread per call" and very primitive synchronization tools are antiquated... then who is he arguing with? Erlang uses actors, Haskell uses sorcery (really, it's fancy; they turn normal-looking threaded code into erlang-ish sliced execution under the covers), Go uses fancy structures along with coroutines, Java uses Executors to implement higher-level work off patterns, and everyone is using Futures and Promises now.
"Simpler" clearly referred to the complexity and the mess created by synchronization in a shared-everything environment, which is where most languages are at with threading (Haskell, Clojure and Erlang are not most languages). This is a valid criticism. You may lose Copy-on-Write, but sharing everything by default, at a low level, does raise the need for complex synchronization, which has its own performance issues.
And this is very good reason to explore and use other concurrency models. This is something that you and the article and the person you are responding to all seem to agree on.
Except that you seem to take any interesting concurrency interface to be threading (no, goroutines are NOT threads) and you draw the strange moral that only the complex problems of threads can be addressed with nice tools and new ideas, but somehow not the problems of other basic concurrency models.
Assuming we have available helpful interfaces for using processes, threads and greenlets - all of which have been produced somewhere - the argument should be about the performance of the foolproofed backends. Preferably based on numbers.
Given that you subsequently make the point that most toolkits are bad, I think my assumption was safe. But let's not lose the plot here, process-level parallelism makes sense in many cases. It's just not a valid replacement for per-process concurrency and it is most definitely not "simpler." Your failure modes become incredibly complex and varied, and that was my only point here.
More modern environments–the ones you should be using unless you have a compelling reason to do otherwise–use shared-state concurrency as a platform for higher level abstractions. But this is still "multicore in-process parallelism" and the underlying model is still threading. Building abstractions on top of in-process parallelism is not rejecting the underlying layer, it's embracing it.
If you are not using modern tools, then yeah correct concurrency is hard and variable degrees of underlying parallelism only exacerbate that. It is also hard to start fires with just flint, steel and tinder. New topic, please.
So I'm not sure what you're taking exception to other than your perception of my tone.
JActor is a high-throughput JAva actor framework capable of delivering up to 200 million messages per second on an i7. It achieves this by operating synchronously whenever it can.
The main innovation here is what I call commandeering, which allows an actor to process the message sent to another actor in the same thread--so long as the target thread is idle.
Consider an event-based server that returns to an event loop whenever it makes a blocking call. At any point in time, most of the active requests will have a call stack depth of zero, and the rest will have just the frames since they were last unblocked. Contrast that with a threaded server where a request occupies a single, deeper stack for its lifetime.
Or consider an actor-based system where an actor has just a few responsibilities and communicates with other actors with messages. Contrast with a threaded system where the server's submodules have to communicate with function calls.
I don't see how stack depth matters much, though. There are way more important considerations.
You knew someone would say it so I will: it matters on RAM-constrained systems. I'm currently working on a CPU with 384K flash (a lot) and 48K RAM (not quite enough).
Many systems in this class use preemptive multi-tasking OSes in a shared, non-virtual memory space. The question how to pre-allocate stack space is tricky. One non-traditional model is to use run-to-completion threads without stack switching so the problem reduces to estimating your overall worst-case stack usage.
Another popular model is main-loop + interrupt handlers. In a lot of ways that's more like an actor model.
My thoughts were:
1) As we enter the massively multicore era, there will be much bigger fish to fry than a little memory overhead. Clock speeds have topped out while transistor counts are still growing exponentially. Exponential means soon application programmers will be tasked with keeping 100s, then 1000s of cores busy. It's gonna be a giant challenge for PL guys, systems guys, and app guys to make that happen. If an improved concurrency model can get us there, then 25% memory overhead one way or the other will be relatively insignificant.
2) The object liveness issue is an implementation detail of one particular VM. There's nothing but CPU cost preventing VM writers from being a little more clever and marking objects unreachable if they are only referenced by dead local vars in otherwise live stack frames. I don't know enough about it to know whether the CPU cost would be prohibitive, but the issue seems more nuanced than just "threads => deep stacks => +25% mem".