Mistakes Node.js Developers Make
airpair.com
airpair.com
This kind of thing would be big enough for me to just recommend to most average developers to just stick with PHP or whatever they are using now. At least the one-thread/process-per-request shared-nothing architecture successfully mitigates the effects of unavoidable developer stupidity (or of "code you've pushed to production without review after more than a couple beers" if you want to put it more 1st person...).
Simply because: 1. bad code will inevitably be written, 2. bad code will inevitably end up, among other things, blocking the event loop, and 3. the application will need to keep working speedly, form the user's perspective, despite having bad code sprinkled in it.
Isn't there any "automagic" way to prevent this from happening with node? ...something like, if a request takes more than XXX ms, than at least start handling new requests in new threads?
This is one of the problems Erlang solves: its internal scheduler keeps things running even if some code ends up in an infinite loop or otherwise misbehaves.
Meanwhile there are actually nice solutions for the problem of handling many concurrent computations with low overhead readily available (e.g. Go).
If a developer can't grasp Node well enough to avoid blocking the event loop, I certainly wouldn't want them trying to handle a thousand connected clients with PHP!
Anyway, to answer your question - within a single node process no, not really. If something on the event loop goes into an infinite loop, timeouts will be of no use, and all other open requests are gonna be toast. This is just a downside to cooperative multitasking.
It is the opinion of many that cooperative multitasking is not the way to go if you want to solve problems which require massive concurrency.
PS: In Racket, you could defend against threads going into infinite loops with custodians.
http://docs.racket-lang.org/more/#%28part._.Terminating_.Con...
Anyway with one nodejs process it's not possible, but you can detect a stuck node and use some load balancing with a cluster of node processes (2 is usually enough), and simply reload a process if it gets stuck (while sending an email to ops of course). You need to use such an architecture to get zero-downtime deployments in any case .
a low priority admin page renders a 3000 row, 30 col table on the server using react. The query and resulting page size are pretty small < 2MB but this takes 5s for react to render.
I didn't expect it to be this slow and can't use client side rendering.
One example: https://github.com/audreyt/node-webworker-threads
Your suggestion is basically to write your own event loop and write any cpu bound task to manually yield to the single-threaded node event loop. Crazy given we have been using multithreaded servers for decades that would do this automatically and generalise for any cpu-bound task.
One can even do rolling upgrades with this for zero downtime releases.
If there is a known workload that will block for a while, one could 1) run the work in a child process, 2) consider using streams to process the data in chunks, or 3) break the work up by chaining smaller operations using "process.nextTick" to put the next operation at the end of the event loop (allows other pending operations to resume before continuing the slow/blocking work).
[1] http://nodejs.org/api/cluster.html [2] https://www.npmjs.com/package/recluster
I sure don't know what kind of developer I am.
Also hot pushes give automatic restart and browser updates out of the box.
Also having optional synchronous api on server, makes a lot of things easier to write.
I like NodeJs, but a lot of aspects will likely prevent it became a mainstream platform.
Node is already a mainstream platform.
Re: 1.1 automatic server restarts (for crashes etc) there's no need for forever / supervisord / nodemon on current Linux: make a .service file for your app and it will automatically restart if it crashes, on all major distros.
[Unit]
Description=My app
[Service]
ExecStart=/var/www/myapp/app.js
Restart=always
User=nobody
Group=nobody
Environment=PATH=/usr/bin:/usr/local/bin
Environment=NODE_ENV=production
[Install]
WantedBy=multi-user.target
From: http://medium.com/@mikemaccana/how-i-deploy-node-apps-on-lin... After=network.target
in the [Unit] section, or you could be in for a bad surprise after a reboot, when your service tries to start before network comes up. Also, I doubt that systemd is available on 'all major' distros.https://github.com/Unitech/pm2
It's significantly more advanced. Create a "process.json" file in your project directory (As you would a nodemon.json), like so:
{
"name" : "someApp"
, "script" : "./index.js"
, "node_args" : "--harmony"
, "log_file" : true
, "merge_logs" : true
, "exec_mode" : "fork_mode"
, "ignoreWatch" : [
"\\.(css|styl|log)$"
, "static/.*"
, "views/.*"
]
}
This will show logs intertwined just like forever, merge logs from seperate forked processes and you can even throw in watch rules like nodemon. pm2 start process.json
Then to make all of your processes always restart do this: pm2 save
pm2 startup centos
Replacing centos with your distro. Couldn't be easier. There are a lot of other features I don't use (such as deployment, which I suspect will be significant for some big websites), but definitely consider dropping forever/supervisord and even nodemon in favor of PM2. They merge pull requests quickly it seems, too.The only issue I've had is that for some reason using the "watch" functionality on a lot of files causes massive CPU overload, or at least it did, it may be fixed by now. If you're going to use the watch functionality (to reload stuff like nodemon), consider whitelisting and not blacklisting files.
Why have a, say, sha1 module instead of a general crypto one? I've seen at least one module that wasn't more than ten lines of actual code. It's packaging JSON and "tests" were far bigger. It was something trivial that should be in a stdlib.
- a terse and frozen API (like "domready" and "xtend") does not end up with the scope creep and bitrot that monolithic frameworks and "standard libraries" tend to carry
- it encourages diversity rather than "this is the one true way"
- it generally leads to less fragmentation overall (and tighter and more robust apps)
- each piece can be versioned, tested and bug-tracked independently
- once you get used to it and start finding modules you like, it can be incredibly easy to prototype and rapidly iterate with existing solutions. My 100+ modules are in a similar domain (graphics) and my efficiency for prototyping has improved because of them.
- it is better for reusability. If you have an algorithm that depends on jQuery or another monolithic framework for just a single function, it is hard to reuse. (ie version issues, bundle size)
It took me a while to come around, but npm has really given me a better appreciation for small modules. :)
Instead, I can find some simple md5 lib (in this case I used one called "js-md5"), and get the same functionality in about 3KB.
It seems a lot of Node projects have gone the small module direction. Given that npm is a package manager that (mostly) works correctly - which is so much harder than it sounds - we can actually use small modules in our app to little detriment.
Yeah, an npm install might take a little bit longer than you'd like, although it's really not that slow. The duplication of modules isn't really an issue in server-side apps, and in webpack apps you can dedupe code pretty easily with webpack.optimize.DedupePlugin + gzipping, so it's not an issue there either.
And I'm not exaggerating about build times. A simple grunt build doing some basic template stuff would take about 20 minutes. The majority of that time was bringing in the ~13,000 files a rather simple static website needed to build. I ended up tossing the idea of independent builds and just made a persistent build machine that symlinked in node_modules. I've got a million lines of C program that takes less time to fully compile and link.
Not sure how your rant about grunt tasks and templating relates to small npm modules. A 20 minute build time sounds like something was vey wrong. My browserify (incremental) build time is < 100 ms which I can handle.
The small module system ends up requiring a ton of files, which is slow. Incremental builds don't really apply to a clean build server where you are basically doing "git clone ... && make". The actual processing isn't my complaint, just the enormous overhead npm's style imposes. I mentioned grunt since just having that plus uglify or so ends up bringing in 13k files or something.
https://www.bouncycastle.org/docs/docs1.5on/index.html
I really don't think the Node ecosystem got this one wrong.
As far as node, you could easily export 20 different hash functions for use, instead of a dedicated sha1 package. Then getting yet another when you want hmac.
The downside is nontrivial. In a repeatable build environment it can add tens of minutes of delays to a build as these tens of thousands of files aren't free. (And that was on a VM by one of the big providers, with top notch bandwidth.) Just even doing a local copy of so many files can take a while.
Why would you be pulling modules in from an Internet based repo on every build?
> Just even doing a local copy of so many files can take a while.
That means your cloud provider has poor file I/O performance, which is not unusual.
Great suggestions on how to profile, but some examples of how to write code that doesn't block could have been given (for example, spawning a child process to do a continuous set of discrete tasks).
Additionally, node's limits should be discussed. If a specifically intense compute task requires a lot of input and will also generate a lot of output (depending on how frequently this task occurs), that's where node starts to really get clobbered and the solution may be very difficult.
How meta.
I don't think this is very beneficial when your project's complexity grows enough to require a non-trivial build. It's nice to have for simple projects, of course.
> Using control flow modules (such as async)
In my experience the async module creates verbose and ugly looking code. In almost all cases promises are a better solution. Of course, even promises are pretty ugly as it's a "hack" to solve a problem at the language level. ES7 proposal includes async/await, maybe that'll finally solve this problem.
>Not using static analysis tools
I'd also recommend checking out TypeScript and tslint for complex Node.js applications.
Even if async/await will only replace .then chains, that already fixes most of the issues. Most async code actually consists of .then chains (or just a single async function call).
Anyway, I agree with your point. Developers need to understand how to use async/await correctly.
I wonder if async/await could be implemented for grouping multiple async operations together..
If doing 'Ctrl-C, node file.js' is too debilitating .....
If you want to forward your logs somewhere remote you can set that up in your syslog config
The whole notion that node.js (or Python's Twisted) makes everything more efficient by putting you in an event loop is just a cop-out for not having something more intuitive like blocking semantics with an evented I/O scheduler.
Or you use Play/Akka and don't think about threads. Or Scala streams and don't think about threads. There's not much thinking about threads going on here. (A major reason why Go just makes me shrug is that I already have its good bits within easy reach on the JVM if I want them, and I don't have its bad bits.)
> If you work in C++, there is no "easy" way, and you have to think about it if you want non-blocking IO.
YMMV, but I don't find boost::thread_group and Boost.Asio terribly hard.
Also, you omitted C#, which has some pretty fantastic asynchronous tools that you don't have to think about at all.
And Akka is perfectly usable in Java.
When management decries usage, it's almost always because they're running a cost/benefits analysis and the costs of using X (including writing language bindings for code, training programmers on the team, hiring new programmers for the team, context switching between multiple languages, and dealing with bugs and corner cases that have been encountered and fixed by other people in more mainstream languages) outweigh its benefits. Many of these costs are invisible to the engineer who originally proposed using X.
But the "repair" I was referring to is that a developer can leave rather than put up with conservative silliness if they so choose (and personally, I do, my last gig was a Scala one and so is the next). Java shops are not so rare as to be irreplaceable and any shop seriously worried about hiring somebody who can work with systems that are by now fairly well-understood isn't going to be a good place for any decent developer's career.
*Anything available in other languages in place of callbacks (promises, generators, async await etc) is also available in node. If you are using callbacks you are just ignorant (or don't want to deviate from out of the box features in which case node is the worst possible thing for you). Nowadays there is not even a performance benefit so the only credible reason to use callbacks is not even there.
console.log('delay is %s', chalk.green(delay));
where as it's much clearer to say
console.log('delay is', chalk.green(delay));
This works in node as of at least v0.10.33 and chrome, but probably most modern browsers.
Shameless plug (co-author here), but if you have any questions about it: feel free to ask.
console.log('delay is ' + chalk.green(delay)));
as it is more consistent with non console.log expressions. For example, just with copy/paste I could replace the above with:
var msg = 'delay is ' + chalk.green(delay); console.log(msg);
But it's really just a matter of preference as long as only one convention is used throughout the same codebase.
console.log(`delay is ${ chalk.green(delay) }`);The prolem with using console.log(object) is that in Node.js the object is not stringified and on browsers the debugger displays the current state of the object, not the state when it was printed. I found that out the hard way..
And when you're using your own logging function you can use a proper logging framework like winston to save the logs to a database and a file in addition to the console (we also send client-side logs to the backend using socket.io for easy debugging).
> console.log('hello', 'world')
hello world
undefined
>
:)That said... I still prefer the format string.
Read the TOC:
1 Not using development tools
2 Blocking the event loop
3 Executing a call back multiple times
4 The Christmas tree of callbacks (Callback light)
5 Creating big monolithic applications
6 Poor logging
7 No tests
8 Not using static analysis tools
9 Zero monitoring or profiling
10 Debugging with console.log
The author talks about trivial or something that has already been written few times.
It would be interesting to see if a post with the title "Top 10 Mistakes C # Developers Make" and comparable content would get just drop the same attention and encouragement. Maybe if it was written by Jon Skeet...