JXcore – A Node.js Distribution with Multi-threading
flippinawesome.org
flippinawesome.org
The touted benefit is marginally increased parallel performance, but even if you buy this (I don't), the best case scenario is this buys you a slightly decreased server bill for the price of closedness, lock-in, and compatibility headaches. If your code shards horizontally like this then you can already trivially parallelize and shard across servers with first-class node core clustering.
I spend most of my waking hours writing all sorts of crazy things in node, and I still can't think of any scenario in which using this makes sense.
> Maybe you don't have a case but there are many scenarios can benefit from multithreading including web hosting.
> The goal is to have a zero compatibility issue and we are almost there with the next beta.
> I believe that the node developers would enjoy benefiting from load aware instance monitored processes instead a trivial multi processing. BTW, still you can combine multithreading with multiple processes in case you want to keep process is active during v8 is GC'ing on one of the threads.
The open source argument holds except in situations where there is a compelling advantage/value. Companies that are dedicated to open source will not be customers, but may not have been paying customers anyway.
MapR is a good example of a closed source technology doing well in the largely open source Hadoop ecosystem. Many companies will not touch MapR, however there are enough companies that really need the features / capabilities / value delivered by MapR that they are willing to pay for it.
Closer to home, there were people who thought Meteor would not take off b/c it dropped npm in favor of a new (open source) package manager. Yet, Meteor provides enough compelling value that it's growing rapidly even with a new package manager.
That said, this product will need to deliver compelling value to a customer segment that's willing to pay, while accepting the fact that many in the node community will not use it. If there are enough of these customers then the company will do fine. However, that's a tradeoff they'll have to evaluate.
As a total aside, it's my opinion that the open source requirement is selectively applied. Mac is clearly not open source, yet BSD is freely available. Yet people choose to use a closed source OS b/c it provides value to them. Same thing with editors and just about everything else. To those who are open source up and down their stack, then kudos for the consistency. For the rest of us, we should at least be honest and own up to the fact that open source is not an absolute requirement, but a selective requirement that's arbitrarily applied. (flame away :)
> You are not alone. remember node.js team members were doing the same in the past but couldn't finish it.
Obviously there are many scenarios this could help.
The problem is solved by running multiple node processes, which is standard deployment for node (i.e. if a machine has 8 cores then 8 node processes are started).
One issue with running multiple threads is that many developers use fail fast as a best practice when building node applications. In other words, uncaught exceptions cause the node process to fail, die and restart. So, it's perfectly acceptable practice to write your application to be written to accept failure as a given (similar to how Netflix uses Simian Army).
That said, how well is each thread isolated from everything else. Does an uncaught exception kill just the thread and its state, or does it kill the process? More specifically, when is the main process killed and what is contained within the thread?
Node applications can be split between stateful servers (ex. chat) and stateless (ex. API for mobile clients).
It's the stateless servers where some developers write fault tolerant / fail fast / fast restart applications. Doing so in a stateful server would be counter productive.
Also, this does not mean to imply that developers are writing sloppy code that fails constantly or that they fail to implement proper error handling. What I was stating is that unhandled exceptions are unexpected, but when they do occur they indicate something is seriously wrong. Importantly, this puts the application's state into an unknown state, which is difficult to recover from. In such situations, a robust approach is to let Node fail, restart fast, and have clean state.
The reasons this is a sound approach are:
1. If the failure is due to a memory leak, then the graphs will highlight said leak clearly
2. The unhandled exception indicates that something is very wrong. A server restart is easily seen in the logs and is a warning that deeper inspection is necessary
3. Recovering state after an unhandled exception is difficult. In a stateless server its better to just restart from a clean state. This assumes the engineers wrote the application to work from a clean state (i.e. after a restart there is no need to recreate state)
4. A fault tolerant architecture is good practice as disks can fail, CPUs can fail, network connections can fail, etc. In a cluster failure is expected and applications are architected to continue operation in the face of failure
(1) request isolation (so most failures in one request can't break other requests) and
(2) a way to catch all exceptions/errors in a single request (and domains don't accomplish this, unless you know what to expect errors from, or wrap everything).
Since node.js doesn't offer those features, it's not even as fault tolerant as PHP was 15 years ago. I'm a huge fan of node.js, but one of the hardest things to do on a large application with a large number of users is to keep an instance of the server from restarting and dropping all the other in-progress requests. If you write your node.js code to be crash-only (like one might do with erlang) your clients are going to have a terrible time.
You may have a 8 cores but configure JXcore to use 64 threads.
> The problem is solved by running multiple node processes, which is standard deployment for node
You can still run multiple node process but 2 threads per each. This will improve the responsiveness of each process by balancing the load exactly on the native side. That means, if anything happens on one of the V8 threads (GC etc). the other one will be handling the load.
> That said, how well is each thread isolated from everything else.
Totally isolated. We already started to update native c,c++ modules for isolated multithreading.
>Does an uncaught exception kill just the thread and its state, or does it kill the process?
On this very beta release, it throws into main thread (when it's uncaught by thread) but coming beta (internal monitoring is implemented) will be optionally resetting the sub thread itself.
You may simply consider each thread as a separate node.js host.
A small diagram on the home page showing x cores * y threads would be interesting / helpful. Such a diagram would also highlight that this adds to the multi-process approach (i.e. I don't have to give up my multi-process approach, but get to add multiple threads to it).
Anyway, the solution sounds interesting.
Other than that they have huge claims about optimizations that the whole node.js community didn't come up with yet and there are just binaries available.
The multithreaded performance optimizations are really interesting, however - but I'd be much more interested in this making it into mainline Node, rather than using this crazy closed-source fork. You're not going to find me using anything closed-source on my servers if I can help it, least of all something like Node which is hackable to its core.
If the performance improvements really mean something, and mean something outside of contrived for() and fib() loops, I hope the Node core teams have a serious talk with these guys about a merge. Node's cluster module is definitely not the final solution for properly utilizing a server's full capacity, and the team knows it; that module has been at level 1 (experimental) for a long time.
from the post "..Last weekend I could finish the prototype and measure the initial performance of the solution.."
If you still understand that it's all finished over a weekend. No way.. I don't think it is easy to develop a LLVM frontend with those features over a weekend.
If you could have some details on LLVM, you wouldn't accuse me on something you miss read.
Because;
There is no such an LLVM JS engine. LLVM has its own IR and that prototype turns the JS codes into a wrapped LLVM IR.. As a result the final code functions at native level, so it is already fast.
For this reason, 20% or more performance difference because of LLVM actually shows how V8 is fast! It doesn't show the prototype is over a weekend or anything else..
Here is where you can download the executable if you want to try it (I found it weird that no link was given in the post): http://jxcore.com/downloads/
But I fail to understand what's the difference between JXcore and native Node clusters. Can you develop this a little bit more?
I don't know about you guys, but I can't rely on some proprietary blob, if not for the security concerns, because doesn't even have an ARM binary, or even an IA32 binary for RedHat derivatives.
What's the core mechanism that permits node's implicit concurrency?
Nope.