A guide to threads in Node.js
blog.logrocket.com
blog.logrocket.com
Server-side JavaScript was a thing long before that. In fact, Netscape used it as early as 1994 [1], predating Java as a backend language. And CommonJS (on what nodejs' modules and many APIs are based) was a community effort towards a common API by 2000s SSJS implementations (helma, v8cgi/teajs, and others).
Apart from that, a nice read for those who need a WebWorker-like API for CPU-bound nodejs tasks.
Funny now there's zero references to it on the internet. I had really liked it back then. It was fully synchronous and had support for threading. Cool stuff:
https://web.archive.org/web/20130615124030/http://silkjs.net...
Edit: Repository seems to be alive! https://github.com/mschwartz/SilkJS/
Sad how neither this, nor xulrunner took off. They simply were ahead of their time. Possibly still are.
In ~1997 you got a version of JScript (very close to JavaScript) in IIS.
Flusspferd—a Spidermonkey-based alternative that was also released in 2009—has quite a nice overview page[0] that's somewhat frozen in time, including "Supported CommonJS Specifications" and "Related projects" sections.
export function runWorker(path: string, cb: WorkerCallback, workerData: object | null = null) { const worker = new Worker(path, { workerData });
worker.on('message', cb.bind(null, null)); worker.on('error', cb);
worker.on('exit', (exitCode) => { if (exitCode === 0) { return null; }
return cb(new Error(`Worker has stopped with code ${exitCode}`));
});
return worker;
}For an article titled "A Guide to Threads in Node.js", I really wish the first example of writing a thread wasn't in TypeScript.
1) Make as few assumptions as possible regarding what the reader knows or has experience with
2) Make code samples as complete and "copy-and-pastable" as possible.
You're assuming that all developers reading this guide will effortlessly translate that TS code to JS.
Using TS creates a possible barrier to understanding, which I'm sure isn't the intent of someone publishing such a guide. Better to simply write some comments or augment the prose in the guide.
I think it would be great if more JS code examples started including type signatures, since they are useful whether or not you program in TS. But yes, there should probably also be a JS-only example in that case and/or a way to toggle them off.
Until TS becomes an official standard supported by ECMAScript out of the box. Were just going along with what feels or looks good and the javascript community has proven that can change from year to year.
I like CS very much myself. While I can see the benefit of having type system (in certain projects), I just can’t get over the aesthetic/syntax issue.
I would much rather do it via Typescripts annotations than JSDocs.
It's not clutter as it was deemed necessary by the writer (me).
If I don't want to use type annotations, I simply don't add them and Typescript does not force you to do so unless you tell it to.
Types aren't coming to ECMAScript anytime soon and with flow and TS it's probably time to get used to the syntax.
No overhead from context switches. Everyone who is running knows how long they have to run, and knows that they are NOT going to get interrupted. No need to worry about locks or how to share data. Want to pass data to another module? Just pass it through well defined interfaces and you darn well know there will never be a read/write conflict, and the next time that module's code runs, it'll have access to that data. (No queue!)
The only exception, and what made it all possible, was the interrupt routines from hardware[1]. Anyone who subscribed to hardware events (entire OS was subscription based, no polling reads ever) had to implement a "thunk pattern" to put data into a receive buffer. The design pattern to do this was the same everywhere in all modules, making code understandable across the entire project.
It is an incredibly freeing paradigm to write in. It becomes so much easier to prove[2] the correctness of code when you can read all the code straight through and not have to ever worry about someone stomping on your data.
It wasn't an RTOS, but even so, it becomes really easy to start providing performance guarantees.
Internal builds had a watchdog timer[3] that would crash the device if it wasn't 'kicked' every so often. Set that to 3ms, start working with the code, and the stack traces tell you instantly who is over their CPU budget. Rewrite code and break it apart into multiple chunks that are scheduled for later execution, repeat until everyone is under their CPU allotment.
For many tasks, single threaded code is nice. Getting rid of preemption is even nicer.
Preemptive multithreading is a compromise. It means that no one thread/process can bring down the system by hogging 100% of resources, but it also creates a huge overhead where important work, work that makes for a better user experience, can (will!) get interrupted for work that honestly doesn't need to be done right now.
The solution to this is just throw so much CPU at the problem that everything gets done in a reasonable amount of time. It has often been noted that "reasonable amount of time" means systems today are less responsive than a 486 running DOS from 1992.
Of course it isn't reasonable to have a modern cooperatively multithreaded consumer OS, no way would the hundreds of processes ran at any one time all cooperate with each other.
But if you ever get a chance to write code on a single threaded cooperative system, go do it. It is a lot of fun.
Now all this meant that going from embedded C to NodeJS wasn't that large of a mental leap! Not having to directly read bytes off the wire was weird (seriously, took a bit of getting use to), but it turns out that "get data, do work, schedule what needs to be done, return early" ends up being the same paradigm at both the top and the bottom of the programming stacks!
[1] All I/O was done using DMA engines, basically a fancy limited programmable piece of hardware that can read and write to all the different pieces of HW hanging off of the main chip, so for example as Bluetooth packets come in, the DMA engine shoves the packets into a buffer and when the buffer is full it raises an interrupt that lets the CPU know that data is waiting. It looks almost exactly like async I/O in any of the modern programming languages, except you have direct access to all that IO being hardware offloaded. Writing to the bare metal rocks.
[2] For a reasonable enough degree of "prove" that software is reliable and doesn't crash from threading issues
[3] A watchdog timer is a physical timer hooked across the power lines of your chip. If it isn't activated every so often (in the industry this is called "kicking the watchdog") it will, in debug builds it does a controlled crash of the CPU, and in retail builds it will reset the entire system. If you've ever had an embedded device reset itself before your very eyes, it is becomes the CPU got locked up and no one kicked the watchdog, so the entire system emergency reset itself.
For the record, I much prefer the term "pet the watchdog" :).
I think the EEs who setup the watchdog were far too grizzled. :D
The reality is, or has become, that the abundance of computational resources has resulted in computers being able to be created in the aether spontaneously, groups of computers even. Making them do one thing in a consistent time span is more important than filling up the theoretical limitation of their allocated resources.
Computer in this context being the elements that allow for computation, as much as an abacus or a professional human has been called a computer. Since many times these compute instances reside solely in the memory of a host machine, somewhere up the chain are the machines made of metal.
A cluster of node processes, of which have their own external memory store and their own database, communicating via REST, is great. The distinction between programming languages being moved into further into irrelevancy, with people that haven't re-learned this crying discrimination, now that their once-useful elitist gatekeeping only serves to remind everyone else how out of touch this person is.
It is fascinating how many computers we actually use now. When factoring in the CDNs over top of our redundant clusters containing single threaded processes with an occasional worker.
It's good to remember that Windows 3.1 was essentially such a system. Apparently, for some reason the paradigm wasn't good enough.
It works much less well when random apps are free to take ownership of the CPU and never release it. :)
This can be addressed by simply giving a higher priority to threads that do UI work. I think my phone works that way, because the responsiveness is pretty amazing.
A single threaded system going all out, with 0 layers of abstraction, can provide a responsive system with a couple hundred megahertz and a dumb display buffer.
Heck at lower resolutions you can get by with under 100mhz and you'll get by just fine.
As an extreme example of this, look at the old video game consoles. Their CPUs ran in the single digit mhz range, but some of them were able to provide frame perfect controls!
This was because they had a code loop that looked like
1. Get inputs 2. run game logic 3. Render graphics 4. Output graphics to screen
Compare that to now days where touch screen latency on some Android phones can be over 100ms! Heck best of class touch on iOS/Android and you'll be at 40-60ms.
Of course the problem with fixed game loops like that is they aren't very flexible. :) My phone needs to do a lot more than render low res graphics to the screen and handle input.
Eventually Android threw enough CPU at the problem that their phones are responsive, but... it took awhile.
Erlang exploits this fact. It is not possible to use iteration, only recursion. This means it is not possible to block the cooperative scheduler unless you have a single incredibly long function that spans several thousands of lines of code. On the other hand calling into C code directly can still bring the cooperative scheduler down which is why Erlang code generally uses a separate process to run C code.
But else, yes, the sent and received data is serialized (when using ipc, else you can also send raw data through streams, and handle the serialization yourself)
ArrayBuffer -> ownership can be transferred around, but the data can only be owned by one thread at a time. The data itself isn't copied. Passing around ArrayBuffers back and forth is good enough for a lot of stuff. In a worker you can then put a node Buffer, TypedArray, or DataView around the ArrayBuffer again, if you want.
I wonder if there are any articles where people use them in threads.
<ducks>