Napa.js: A multi-threaded JavaScript runtime
github.com
github.com
Now if we started talking about "computeless" architecture I'll be confused. (Though maybe that'll be the trendy name for serverless data sources/sinks in a few years...)
I am pretty sure in English serverless means no server and "unprovisioned/unconfigurable" machines means you didn't provision them and you cannot configure them. Even in analogical sense this makes no sense. Something i could relate to is something like "Pay as you use" or "configurationless servers".
But that is just me, and if you think it is ok to randomly change the meaning of words that means me personally and randomly don't need to accept your new meaning (not giving out, just trying to explain my rationale)
Downvote all you want, but please do point out where i am wrong.
Language is all about context. The meaning of words changes depending on it, even in plain English settings.
Wide spread terminology suffers from the evolutionary pressures of marketing. Only the catchiest, most marketable terms propagate.
Because speech (and writing) can be figurative, not just literal. And because the term has reached wide adoption (at least, in the subset interested in discussing such things) and so not using the term makes conversation difficult and litigating the issue every time it's discussed adds no value to the discussion whatsoever.
For what it’s worth, I don’t think calling them “configuration-less” or “pay as you go” servers is any more accurate. You’re really just buying processing time.
So eventually people picked one. Today, the most common are "serverless architecture", "FaaS", or simply "Lambda" (borrowing from AWS).
You don't have to do anything. But it's simply a fact that many people know what you're talking about if you say the word "serverless". And that's what language is, a (kinda) agreed upon set of words which let you communicate with other people. If everyone but you understands a word, and you are crusading that they change it to something else, what is the point?
If you're interested, the concept of "prescriptivism" may be enlightening.
Lots of discussions have been had on how many people think "the cloud" is a sorta-magical thing, which "is just there". Just more recently some interesting aspects of "using the cloud" have been more thoroughly discussed (e.g. the jurisdiction it's hosted in, data breaches, etc). If the concept was described in a less abstract way, would these discussions have happened sooner? Later? Would it have become less of a "buzzword" amongst executives?
So, is "cloud" really a better term than "serverless"?
Could there have been a better word than serverless? Probably, but that is the one that is currently used for that general kind of architecture. I would have called that PaaS before, and sometimes still do.
In reality, words are often used in ways that don't necessarily meet the dictionary definition in the strictest sense.
For example, I complained to my local advertising authority that mobile providers are using the word "unlimited" to mean "limited by our fair usage policy" and I was told that this is fine as long as 95% (I don't remember the exact percentage, maybe it was 99) of customers will never reach the limit so its effectively unlimited. That's not really what the english language word means, but hey... that's life. Same thing applies here: words are recycled to have different meanings.
There are situations when I think its best to just go with the flow of how people are commonly using words but simply accepting what others say in a blanket sense is not right either. In the case of "serverless" it doesnt really matter much to me, but if you think of the word "gyp" for example thats something many people have had to make a conceited effort to stop using. So in some cases, with effort we can improve the language we use and not always be swimming against the tide.
So it does allocate virtual infrastructure?
If you were to go down deeper, a server is just an electrical machine which shuffles eletrons around. So, you could say "there's no point in talking about 'servers', if we're just using transistors when you really think about it".
But it would convey no useful information if someone asked you "what are you using to run your service?" and you replied "well, I just move electrons around", would it?
As far as managing goes, most of the server are not managed by you anyway. But yeah i get it, everyone likes it so...
In that scenario, not having to even consider if the error you're getting is because your server is restarting/out-of-memory/needing-update/broken-by-a-coworker like 99% of the time is as close to "serverless" as it gets, in terms of day-to-day work activities and worries.
Project is a React client and express backend. It's built and tested dockerized. We use testcafe and Chrome headless, so more memory is always useful for parallel builds.
Likewise, the distinguising factor between multi-threading and multi-processing is shared memory, i.e. again, the speed of communication.
Multi-threading does well for some problems, but often, multi-machining or multi-processing is sufficient, which is why so many runtimes don't really do multi-threading: Node.JS, Python, Ruby.
The benefit to programs that don't use threading and use event loop and shared nothing multi-process is that they don't have the overhead when things are maxed out.
This is why virtually every high performance server (nginx, redis, memcached, etc) is written this way and things like varnish (thread per request) are multiples or orders of magnitude slower.
Funny people criticizing nodejs for using the same architecture that all the best-in-class products use.
This might have been what you meant when you said the runtime isn't really multi-threaded - but since the CPython ecosystem rests so heavily on C code, in practice multithreading is a good solution a lot of the time.
Node.JS, Python, and Ruby are in C/C++ so you can do multithreading in all three, if you are willing to write native modules.
Better than C/C++ sometimes.
Imagine buying a new desktop computer, not the most expensive but a good performing one, and setting it up at home to serve some kind of cloud services. With its cpu fully utilized 24x7, I bet buying equivalent compute at AWS would be crazy expensive per month.
Of course there are many reasons most people don’t use desktop systems as home servers, but I bet there are a few scenarios where it could payoff.
One example might be a bootstrapped startup tight on cash flow, with cpu biased workload, so ISP bandwidth and local disk throughput didn’t bottleneck before the compute did. And it couldn’t be a mission critical service, something where some maintenance windows wouldn’t kill your reputation.
Finally you’d need a way to really minimize admin/devops costs. That kind of work doesn’t take many hours to kill your savings not to mention opportunity costs.
Of course, that doesn't include storage and bandwidth.
I mean, sure. I've got a 5 year old laptop that will outperform the t2.micro I'm pay $8.35/month for. But I don't trust my home internet to be stable or fast enough. Not to mention that my primary usage is an IRC bouncer, so I need it to not be on my home internet connection so some script kiddie doesn't DDoS me after I ban them from a channel because they were spamming racial slurs. Yes, that has actually happened.
Now I would not host my main customer site there. But dev servers? Beta servers? QA servers? Hey why not.. save some massive bills.
I average probably a single 5 minute hiccup each month. That's 99.988% uptime. For someone wanting to run their WoW guild's voice chat server, or just a toy server, or a development/staging environment, that's plenty.
But I mean, my home internet is only 35 mbps anyways via Frontier FIOS. I can get 150 mbps through Comcast, but I refuse to give that company a penny of my money. In either case, I'm not going to be running any major production servers at home anyways.
used ThinkPad W540 with same config goes for ~1.5 months of your AWS rent.
Laptops are surprisingly good as little dev servers. In fact, you can find ones with broken screens for even cheaper, which is fantastic!
Few hundred bucks can get you a nice i5 or i7 processor
https://news.ycombinator.com/item?id=4063929
I wrote this up: https://news.ycombinator.com/item?id=15499629
“I believe this to be a basis for designing very large distributed programs. The “nodes” need to be organized: given a communication protocol, told how to connect to each other.” ttps://www.americaninno.com/boston/node-js-interview-4-questions-with-creator-ryan-dahl/
I don't understand why people keep insisting that the lack of multi threading support is a Javascript problem when there are better and more scalable ways of using your machine resources.
Not ideal, but not too bad either.
I completely agree with what you’re saying for those other problems you describe. But for the typical node.js webapp, doing the one-container-per-core is perfectly fine.
My app should be stateless and scale by replication. This is the most efficient solution in every possible way.
Containers don't cost meaningfully more than threads unless you create expensive unique resources for each one.
This is way more overhead than threads which can share all of those resources.
Shared process memory isn't the easy memory-consumption win it sounds like, locking is hard to get right, potentially very destructive to the parallel performance that was the point of the whole exercise, and marries you to a single physical box.
Even if you want to take advantage of shared-address-space shared memory you probably want to do it in a more principled way than fork()
One-copy-per-thread and share-by-communicating both give you braindead simple scaling without dealing with that.
What if you have a large shared in-memory data structure that you want to update with lots of irregular translations in parallel? Like many graph problems? How are you going to do that with multiple containers? As in industry we just don't understand how to distribute that kind of problem effectively.
One of the reasons Node is successful is the simplicity of single threaded code. Way easier to reason, I would question the usage of Node if you are doing something CPU bound with it. You can use golang or C# with tasks for that.
If you really want to do something creative with the shared memory, I guess you could do that in a "native module" written in c++ or even Rust[1].
I'm not saying that it's not doable with JS, it's just that it's already been done (as in, has a solution that works).
Right tool for the job.
> What would happen if we made everyone program enterprise CRUD applications in C++ from scratch
I don't think that's a good analogy. A better analogy would be if we started writing everything in JS/Ruby/Python instead of using lower-level languages where performance matters. Except, this is regularly done with great success by many, many companies, so I don't think that helps your point. Sure, you may have to eventually port it to a more performant platform when you hit massive scale, but that point may also never come.
Most cases, when you need to parallelize and multiprocess doesn't cut it, you'll likely need to go to the C++ module route anyway.
I think if there is a sensible use-case for parallel computing in JS, it would be good to have. However, trying to make a solution before we have a (clear) problem is foolish.
I'm not saying there isn't already a use-case, but I haven't seen one that isn't already covered by languages better suited to solving those problems (e.g. Rust).
Edit to give a different example: parallel computing in JS is like trying to write a web framework in Rust. Sure, you can do it, but Node is already better suited to doing that. At best, you're making a worse version of something that already exists.
I agree with you, but my point is it's not black and white. For a sufficiently small or simple project, it might make sense to write a web backend in Rust, or do parallel computing in JS, if the cost of learning a new platform outweighs the cost of using the "wrong tool".
In most circumstances, yes, you probably shouldn't use Node.js for parallel computing tasks, just like you shouldn't use C++ for web development, but for some use cases it might be useful. And maybe those use cases don't exist (I don't have much experience in this area, so I don't know), but I just don't like when blanket statements like "use the right tool for the job" dismiss the work other people have done. Surely if Microsoft created this, they have a use case in mind for it?
You make it sound like it was difficult to learn. Underneath, C++, Java, Pascal, C#, Javascript and Python, have many similarities and jumping from one of those languages to another in the list is very easy; compared, for example, to something like jumping from any of those languages to Forth, PROLOG, SQL, ML, Haskell, or Lisp.
Some of them are also really similar syntactically, for example this group: [C, C++, Java, C#]; or this other group: [Pascal, Algol, Go], so even the syntax doesn't get in the way when jumping from one to other.
Thus, usually, software engineers do know more than one language and they apply what better suits the program.
Think about something like Delaunay triangulation or mesh refinement. These are critical path bottlenecks for a great many applications and in practice very parallel, but they're irregular so we cannot easily distribute the data structure. The best results we have are for shared memory thread models. We don't know how to do it any other way!
Things like greenlet and gevent (and likely napa.js) are band-aids over the underlying problem.
Redis also might not be the best choice if thats your primary use case...but still.
Then we got a very fast JIT, and suddenly you could do reasonable compute heavy stuff very fast, and then it became viable to also write the server side in JS, because of programmer efficiency and library reuse and other reasons.
The "right tool for the job" can seriously change when tools improve and develop, and just because there already are other tools for the same job should not stop anybody from trying.
I can not think of a better example for that than JavaScript.
Redis probably isn't a great example here. I've worked on projects where a single Redis instance was not enough (would easily peg its single CPU to 100% and have query latency in the multi-second range). In the end, sharding the data among several Redis instances was successful, but also brought its own problems. The ideal is that we just have languages, runtimes, data stores, etc. that abstract these details away from us so we can focus on our application logic, not on how to make it faster.
This seems like a bit of a FUD.
With multiple threads and shared data, you don't necessarily have to share all the data structures with all other data structures and all the threads. You can setup your things such that minimum or nothing is shared. That's (also) what access control and immutaibility is for in programming languages, apart from other features.
Of course, different languages support these features in different ways, I don't want to get into the specifics, but in pretty much all mainstream languages you can create a similar share-nothing or share-almost-nothing design and it's not even hard, it might even be easier.
I really don't understand modern web/JS developers. They seem to ignore traditional solutions and/or proclaim them as evil, and then they go on to employ a 'new' solution that is 3× as complex, performs 5× worse and requires 10× as many dependencies/tools/frameworks/etc. Why? I suspect there's a LOT of largely irrational fear of concepts and languages that are unfamiliar. "Fear driven developement" in fashionable lingo.
TL;DR you don't need to be scared of threads, you just need to be scared of threading architectures that share too much.
It is, perhaps, because a significant amount of Node.js developers came from front-end-only development, thus unfamiliar with the traditional approaches (in this case, using threads). An example is the many cases in which a document store as MongoDB is (wrongly) used for data that is mostly relational.
Simply put, they never were taught the traditional approaches first.
Use C++.
If you ever decide to scale and distribute it for real save the state in a database and orchestrate containers. For this solution i would recommend using Node.js or Go.
My point was that this isn't true:
> You can get rid of the need to multi threading by deploying more containers in the same machine or via orchestration.
If you have a large shared memory data structure and irregular updates then this approach won't work, no matter what language you are using. If you say use a database instead, well then the database just has to solve exactly the same problem, and they'll use shared memory parallelism as well.
You can punt the problem further down the stack, but some, somewhere at some point is going to need to solve the problem, and they're going to use shared memory parallelism to do it.
It works for specific types of work, but not necessarily as a low cost abstraction unless you pre-thread. In other words, it works well for some cases where threads are used, and horribly for others.
1. Setting up containers and the communication between them is really complex - it also only makes sense for a deployed, long-running, server process. 2. Communication between containers is very expensive. You're never going to beat an in-process pointer to shared memory.
Web workers would make a lot more sense for most javascript programs.
2 - That is only true if the your bottleneck is on communication somehow. Usually it is not.
And a huge advantage: You will write stateless and easy to scale apps that do not care how the machine resources are being handled.
It also means accounting for each system failing in your own app, retries with exponential backoff, timeouts, logging errors, circuit breakers, plus all the exotic ways a network layer can fail - if your RPC protocol has arbitrary limits (message size, timeouts, etc)
You will probably want to use kubernetes with istio, not raw docker. All very do-able, but definitely not simple.
I agree that services make sense, but there's a level between single-threaded and micro-services where having concurrency within your application is useful.
https://github.com/MicrosoftArchive/redis/issues/556 (Jun-Sep 2017)
>Why do MS always do this??? Start something, announce it aloud "we are now open source, we are now this and that blah blah" then quietly do a 360 and moonwalk away
> do a 360 and moonwalk away
https://en.wikipedia.org/wiki/Moonwalk_%28dance%29
>moves backwards while seemingly walking forwards
It's actually a near-perfect analogy here, where actions speak louder than words re: Microsoft's commitment to maintaining adaptations of existing open source projects. I think the example provided is enough to trigger a careful analysis before jumping in, and I would love to see counter-examples to help balance the evaluation.
(Please note: very specifically requesting counter-examples of Microsoft-official, intended for production, open source repositories demonstrating long-term maintenance of tweaks of/dependencies on established open source projects that for whatever reason were never up-streamed.)
'Why do they call it the Xbox 360? Because when you see it you do a 360 and walk away.'
They built this to scratch an itch, then they opened the source, but that's still not enough?
If the maintainers of a project don't manage the project properly and other people think they can do a better job, Open Source gives you the right to fork. And this social contract is an incentive for the maintainers to do a good job in order to not lose control.
On the other hand the maintainers have no contractual obligation to keep supporting the project.
When maintainers stop supporting the project, the code is still there to be picked up if there's enough interest. That's powerful, but also distributes the responsibility to all interested parties.
If no new maintainers show up to fork the project, then maybe it's OK for the project to die.
When Windows Phone 7 was the new thing, I went to one of the Microsoft dev camps for it. I vividly remember someone in the audience, standing in the walkway berating the guy on stage talking about Windows Phone 7 development because he had been focused on some Microsoft technology (Silverlight? WPF? I can't remember) and they had basically relegated it to the past by switching over to this new framework. At the time, I was kind of just mind blown, but in retrospect I can see his disappointment. If it were open source, either another entity could champion it or it would still go in the heap of bygone frameworks, but at least then it stood a chance.
And I'm fine with Microsoft playing around with something, open sourcing it, and then they decide there is something else out there they want to work with. Mostly in cases like this where it isn't a "product" they're attempting to market. Seems to be an okay process and perhaps someone can look at what they left in their wake to gain some kind of insight from it.
> They built this to scratch an itch, then they opened the source, but that's still not enough?
I'm honestly not sure where this question is coming from. The issue I linked is on a port of Redis to Windows done back when Microsoft Open Technologies was temporarily spun out as a subsidiary, and is one of many there reflecting frustration at Microsoft's mismanagement of the project, especially lack of communication regarding long-term support. It serves as a recent example of how Microsoft drops a low-priority open source project.
Microsoft is building a strong track record with open source projects that they completely control, but less so when they don't. I believe this is relevant here because of Napa.js's tight coupling with Node.js.
Microsoft Open Technologies had some successes here too: off the top of my head NodeJS and Git both have gotten much better on Windows as a result of work that Microsoft Open Technologies started in forks and eventually got merged upstream. The NAPI work to support both V8 and ChakraCore continues at a reasonable pace precisely because there is upstream engagement, upstream merging, and NodeJS-ChakraCore is less of a full fork more like a distro-specific build flag.
Plus, the priority on things like Redis for Windows shifted with the Windows Subsystem for Linux; there was less need to get open source projects to treat Windows as a supported distro directly when Windows can piggy back off of Ubuntu (or SUSE or Fedora) distro work.
None of that solves lack of communication about the forks left on the vine, and it is a shame there isn't strong redis support on Windows, but everyone has priorities, including antirez, and even if those priorities aren't communicated fully, they seem to at least be somewhat transparent to this humble developer from the outside of either work.
In my mind, it will have to attract a strong enough community to survive without Microsoft's financial support which will probably shut down in a year or two. It's interesting because if it does hit critical mass that increases the likelihood that it will continue to be funded.
In the NodeJS world, sometimes "long-term" means 9 months it feels like. In open source, sometimes "support" means "file a PR if you care so much about that bug" or "fork it yourself". Are you perhaps expecting a longer term or more support because it is Microsoft behind it?
It looks actively maintained right now, but its published semver is 0.x. There are a lot of 0.x libraries in active use in NodeJS, but it's still semver-obvious grain of salt from the maintainers to keep in mind when considering it for support.
The README tells us that it was directly built to support a production need in Bing today. That seems like as big of a vote of support confidence as you might get from an open source project that a team has vested interest in it.
On the other hand, Bing isn't inside Microsoft's developer division, so they have fewer vested interests in supporting outside developers long term. It's also possible that at some point in the future they get internally sold on a developer division or Windows division or Azure alternative that Microsoft has financial or marketing reasons to commercialize.
If I had a production need for something like Napa.js I don't see any particular red flags to avoid it, it looks like it should be easy enough to migrate to something else down the road if necessary, and would probably consider it.
https://en.wikipedia.org/wiki/River_Trail_(JavaScript_engine...
Another attempt (although a different approach) is here : https://github.com/tc39/ecmascript_sharedmem
Hopefully, Microsoft will learn from the errors of others. Interestingly, napa is build on V8, not on MS own javascript engine.
https://webkit.org/blog/7846/concurrent-javascript-it-can-wo...
1. It is exposed as a nodejs module, but some other modules may be at conflict because they may wrongly assume that they are the only thread running. E.g., some modules may be using global variables in C++.
2. Still, it doesn't support fast immutable communication (structural sharing) to other threads through shared memory.
Does Java offer cross-thread shared memory? Go, Haskell, Python, etc.?
Python supports threading but is not very useful due to the infamous GIL. However, multiprocessing is usually a decent alternative and it supports shared memory on forking (copy on write, though). Also you can use mmap easily for IPC.
For anything else, IMHO the use-case will be so narrow, that I wonder how MS will justify the development resources for maintaining it.
As far as I can see, cluster workers are all uniform (no task-specific pools) and cluster has no broadcast (just `worker.send(message)`.)
"Use Isolates for secure, concurrent apps. Spawn an isolate to run Dart functions and libraries in an isolated heap, and take advantage of multiple CPU cores." -- https://dart-lang.github.io/server/server.html
Then, this has a lot of compatibility issues, and release-wise it will be hell. I don't think this is a good idea.
For complex objects, usually JSON is used thus marshall/unmarshall is needed. But for objects like UTF-8 string or ArrayBuffer, the same layout is used across JS and C++, thus at almost no cost.
Another thing is that between addon-modules, they can pass pointer of native structures (like Buffer) using 2 uint32 through JS. In this case, JS works as a binding language.
I would assume this would benefit desktop JS most, so it'd undoubtedly help Electron apps.
EDIT : There's a better explanation at https://github.com/Microsoft/napajs/wiki/Why-Napa.js
Threads are great also when you need to do CPU intensive calculations such as processing an audio file or number crunching. That's what NapaJS seems to be geared towards in fact.
As people pushed the limits of Python's performance and were only met with answers like "use C for fast code, the GIL is probably never going to be removed", Go and Swift emerged, clearly intending to provide a Python-like development experience with performance characteristics more similar to Java, including the ability to run real threads.
Other runtime developers should pay attention and be careful not to fall into the same trap.
Seriously... when publishing a project like this, the authors ought to give a least a bit of insight into their motivation. Their examples include calculating Fibonocci in parallel, not exactly breakthrough.
Atom's selling point to me was always its easy hackability, and that's where Electron tremendously wins: doing hacks, software meant to be set up quick'n'dirty, personal experiments, stuff done because "why not". I'm glad that Electron (and Atom too) exists because of that. However, that's definitely not a reason to write any serious tools in it. Using any Electron app with my 8GB RAM is a nightmare, and most of them look so out of place that (from the UX perspective) I don't know why they were put outside of the browser in the first place.
[eidt] Oh, and the main thing - VS Code and Atom definitely don't share their code with any Web version running in browser. Doesn't that make the original argument moot?
At least it's fast.
EDIT: found one that doesn't, but it clearly loads some weird custom stylesheets, since it doesn't look like anything else.
However, that's a really bad argument when it comes to software development of proper tools that need to be developed and maintained and actually used by people.
Why I use Object Pascal | https://news.ycombinator.com/item?id=15490345 (Oct 2017,250+points,215+comments)
But for node cluster's case, if we need to load a data set for a worker to serve request, it has to be loaded on each process. While the multi-threaded model will have only one copy.
Edit: I would question usage of Node if everything you are doing is CPU bound. I think people started down-voting me without going through docs, I mean just read this documentation page and tell me this is right https://github.com/Microsoft/napajs/blob/master/docs/api/mem...
> As it evolves, we find it useful to complement Node.js in CPU-bound tasks, with the capability of executing JavaScript in multiple V8 isolates and communicating between them.
Not everything MS did was evil.
For example, let's say your Node application's requirements changed, and now you have to transform a set of five JSON objects to one JSON object, with some big arrays, indexing, and sorting required. The operation takes 10-100ms of cpu time, which isn't nuts, but with thousands of requests per second would normally be a catastrophe. If you can spawn a thread in this case you can save the day.