I don't understand why people keep insisting that the lack of multi threading support is a Javascript problem when there are better and more scalable ways of using your machine resources.
I don't understand why people keep insisting that the lack of multi threading support is a Javascript problem when there are better and more scalable ways of using your machine resources.
What if you have a large shared in-memory data structure that you want to update with lots of irregular translations in parallel? Like many graph problems? How are you going to do that with multiple containers? As in industry we just don't understand how to distribute that kind of problem effectively.
One of the reasons Node is successful is the simplicity of single threaded code. Way easier to reason, I would question the usage of Node if you are doing something CPU bound with it. You can use golang or C# with tasks for that.
If you really want to do something creative with the shared memory, I guess you could do that in a "native module" written in c++ or even Rust[1].
I'm not saying that it's not doable with JS, it's just that it's already been done (as in, has a solution that works).
Right tool for the job.
> What would happen if we made everyone program enterprise CRUD applications in C++ from scratch
I don't think that's a good analogy. A better analogy would be if we started writing everything in JS/Ruby/Python instead of using lower-level languages where performance matters. Except, this is regularly done with great success by many, many companies, so I don't think that helps your point. Sure, you may have to eventually port it to a more performant platform when you hit massive scale, but that point may also never come.
Most cases, when you need to parallelize and multiprocess doesn't cut it, you'll likely need to go to the C++ module route anyway.
I think if there is a sensible use-case for parallel computing in JS, it would be good to have. However, trying to make a solution before we have a (clear) problem is foolish.
I'm not saying there isn't already a use-case, but I haven't seen one that isn't already covered by languages better suited to solving those problems (e.g. Rust).
Edit to give a different example: parallel computing in JS is like trying to write a web framework in Rust. Sure, you can do it, but Node is already better suited to doing that. At best, you're making a worse version of something that already exists.
I agree with you, but my point is it's not black and white. For a sufficiently small or simple project, it might make sense to write a web backend in Rust, or do parallel computing in JS, if the cost of learning a new platform outweighs the cost of using the "wrong tool".
In most circumstances, yes, you probably shouldn't use Node.js for parallel computing tasks, just like you shouldn't use C++ for web development, but for some use cases it might be useful. And maybe those use cases don't exist (I don't have much experience in this area, so I don't know), but I just don't like when blanket statements like "use the right tool for the job" dismiss the work other people have done. Surely if Microsoft created this, they have a use case in mind for it?
You make it sound like it was difficult to learn. Underneath, C++, Java, Pascal, C#, Javascript and Python, have many similarities and jumping from one of those languages to another in the list is very easy; compared, for example, to something like jumping from any of those languages to Forth, PROLOG, SQL, ML, Haskell, or Lisp.
Some of them are also really similar syntactically, for example this group: [C, C++, Java, C#]; or this other group: [Pascal, Algol, Go], so even the syntax doesn't get in the way when jumping from one to other.
Thus, usually, software engineers do know more than one language and they apply what better suits the program.
Think about something like Delaunay triangulation or mesh refinement. These are critical path bottlenecks for a great many applications and in practice very parallel, but they're irregular so we cannot easily distribute the data structure. The best results we have are for shared memory thread models. We don't know how to do it any other way!
Things like greenlet and gevent (and likely napa.js) are band-aids over the underlying problem.
Redis also might not be the best choice if thats your primary use case...but still.
Then we got a very fast JIT, and suddenly you could do reasonable compute heavy stuff very fast, and then it became viable to also write the server side in JS, because of programmer efficiency and library reuse and other reasons.
The "right tool for the job" can seriously change when tools improve and develop, and just because there already are other tools for the same job should not stop anybody from trying.
I can not think of a better example for that than JavaScript.
Redis probably isn't a great example here. I've worked on projects where a single Redis instance was not enough (would easily peg its single CPU to 100% and have query latency in the multi-second range). In the end, sharding the data among several Redis instances was successful, but also brought its own problems. The ideal is that we just have languages, runtimes, data stores, etc. that abstract these details away from us so we can focus on our application logic, not on how to make it faster.
This seems like a bit of a FUD.
With multiple threads and shared data, you don't necessarily have to share all the data structures with all other data structures and all the threads. You can setup your things such that minimum or nothing is shared. That's (also) what access control and immutaibility is for in programming languages, apart from other features.
Of course, different languages support these features in different ways, I don't want to get into the specifics, but in pretty much all mainstream languages you can create a similar share-nothing or share-almost-nothing design and it's not even hard, it might even be easier.
I really don't understand modern web/JS developers. They seem to ignore traditional solutions and/or proclaim them as evil, and then they go on to employ a 'new' solution that is 3× as complex, performs 5× worse and requires 10× as many dependencies/tools/frameworks/etc. Why? I suspect there's a LOT of largely irrational fear of concepts and languages that are unfamiliar. "Fear driven developement" in fashionable lingo.
TL;DR you don't need to be scared of threads, you just need to be scared of threading architectures that share too much.
It is, perhaps, because a significant amount of Node.js developers came from front-end-only development, thus unfamiliar with the traditional approaches (in this case, using threads). An example is the many cases in which a document store as MongoDB is (wrongly) used for data that is mostly relational.
Simply put, they never were taught the traditional approaches first.
Use C++.
If you ever decide to scale and distribute it for real save the state in a database and orchestrate containers. For this solution i would recommend using Node.js or Go.
My point was that this isn't true:
> You can get rid of the need to multi threading by deploying more containers in the same machine or via orchestration.
If you have a large shared memory data structure and irregular updates then this approach won't work, no matter what language you are using. If you say use a database instead, well then the database just has to solve exactly the same problem, and they'll use shared memory parallelism as well.
You can punt the problem further down the stack, but some, somewhere at some point is going to need to solve the problem, and they're going to use shared memory parallelism to do it.
Not ideal, but not too bad either.
I completely agree with what you’re saying for those other problems you describe. But for the typical node.js webapp, doing the one-container-per-core is perfectly fine.
My app should be stateless and scale by replication. This is the most efficient solution in every possible way.
Containers don't cost meaningfully more than threads unless you create expensive unique resources for each one.
This is way more overhead than threads which can share all of those resources.
Shared process memory isn't the easy memory-consumption win it sounds like, locking is hard to get right, potentially very destructive to the parallel performance that was the point of the whole exercise, and marries you to a single physical box.
Even if you want to take advantage of shared-address-space shared memory you probably want to do it in a more principled way than fork()
One-copy-per-thread and share-by-communicating both give you braindead simple scaling without dealing with that.
It works for specific types of work, but not necessarily as a low cost abstraction unless you pre-thread. In other words, it works well for some cases where threads are used, and horribly for others.
1. Setting up containers and the communication between them is really complex - it also only makes sense for a deployed, long-running, server process. 2. Communication between containers is very expensive. You're never going to beat an in-process pointer to shared memory.
Web workers would make a lot more sense for most javascript programs.
2 - That is only true if the your bottleneck is on communication somehow. Usually it is not.
And a huge advantage: You will write stateless and easy to scale apps that do not care how the machine resources are being handled.
It also means accounting for each system failing in your own app, retries with exponential backoff, timeouts, logging errors, circuit breakers, plus all the exotic ways a network layer can fail - if your RPC protocol has arbitrary limits (message size, timeouts, etc)
You will probably want to use kubernetes with istio, not raw docker. All very do-able, but definitely not simple.
I agree that services make sense, but there's a level between single-threaded and micro-services where having concurrency within your application is useful.