Node.js - A giant step backwards?
fennb.com
fennb.com
I hate tabloidy headlines with question marks. As in "Queen Elizabeth: Is She a Transvestite?", "The Moon: Is It Made of Cheese?" or "Linkbait: Will It Ever End?".
There really should be language support for this sort of thing (like coroutines) so these sort of cascading changes don't need to happen.
I'm all for using coroutines to solve this problem. That's the approach taken by my Celluloid::IO library:
https://github.com/tarcieri/celluloid-io
Unfortunately Ryan Dahl is adamantly opposed to coroutines so that's not going to happen in Node any time soon.
What would make Node more attractive is if it supported copy-on-write multithreading and gave me a way to cheat and use asynchronous I/O (like a wait(myFunctionThatTakesACallbackOrDeferred) function)
https://github.com/athoune/em-scenario
Note: I still think this approach sucks.
V8 provides a really awesome shared-nothing multithreading scheme via web workers. It's just nobody uses them.
"Node does not modify the JavaScript runtime. This is for the ECMA committee and V8 to decide. If they add coroutines (which they won't) then we will include it. If they add generators we will include them."
See Tim's thread, My Humble Coroutine Proposal (https://groups.google.com/forum/#!topic/nodejs/HJOyNMKLgB8). Warning: long.
So, you don't like javascript? Don't use it.
But you won't beat the control over program state offered by javascript -- not until computers understand their programmers well enough to reason about program state. We need better computers, better computer science, and better programmers -- that's all. Then we can replace NodeJS with something better.
Until then, NodeJS is almost certainly the most tasteful solution to the most common problems. I hope its replacement meets so high a standard.
Edit: I guess a hopeful idea is that there should be no reason something like call/cc could not be added to Node.js. In that case, the extensive library of non-blocking functionally will be very handy since you could build a sane continuation based interface on top of it and escape from callback purgatory.
Not that it can't be used under the hood of course.
However, monads don't solve this problem - they cause it, since their primary concern is correctness and not converting between monadic and non-monadic code. If you have a pure function and need to convert it to a monadic action there will be lots of collateral damage as functions that interacted with the old function have to be converted to monadic style.
You don't have to convert them, but you do need a way to write code using those pure functions "inside" the monad which I don't think is easy with this model.
You should be able to just chain your actions together, with a wrapper function to turn pure functions into actions. I don't see why you would need to convert the actual pure functions to monadic style.
I guess what I had in my head when I made the comment was the convenience functions or the syntactic sugar around a monadic solution. Seems like the node guys are playing with things like this but all the solutions I've seen seem so ugly.
//turn a function into a monadic action
function lift(f){ return function(/**/){
var d = new Deferred(); //Im using the Dojo API
d.resolve( f.call(this, arguments) );
return d.promise;
};};
var f = function(x){ return x+1}
lift(f)(0)
.then(function(y){ ... })
The problem is the opposite direction: pure code using monadic code and pure code turning into monadic codeIf you have something like
var x = f1( f2( f3() ) )
And f2 becomes a monadic action you have to rewrite this bit as var xPromise = f2( f3() )
.then(function(f2result){
return f1(f2result);
});(isn't the opposite direction the feature of doing things this way? That you can't call monadic code from pure code is a good thing)
Absolutely, I've had it happen in "big" (client-side) JS projects.
Only way I've found so far to handle this is to make anything which might ever have any reason to become asynchronous (so anything but helper functions) take callbacks or return deferred objects. Always.
But then an other issue arises: for various reasons, callbacks-based code which works synchronously may fail asynchronously, and the other way around. And then it starts getting real fun as you still have to go through all your (supposedly) carefully constructed callbacks-based code to find out what reason it would have to fail (alternatively, you create both sync and async tests for all pieces of callbacks-based code)
The example "asynchronous get" becomes something like:
local myObject = query{ id=3244 };
The query function can be written to handle a cache lookup, a database query that gets stored in the cache, and any other logging you want, because Lua has coroutines. All I/O that gets sent through the "ngx" query object (which can connect to local or remote ports) yields control to the main loop."query" is an example of a function you could create; its implementation (with a cache lookup) could look something like (yes, I use CouchDB...):
function query(t)
local result = ngx.location.capture( "/cacheserver/id:"..t.id );
if #result == 0 then
result = ngx.location.capture( "/couchdb/usertable/"..t.id );
end
return result
end
I've heard reports of 50k+ connections/second on a VPS running Nginx, LuaJit, and the LuaNginxModule, and on my low-end VPS it easily handles 2000+ connections per second (with CouchDB queries) with no more than 250ms latency. Actually, that was as many connections I could send at it, so it may be able to handle a lot more.Node sucks for general web apps because you have to program everything asynchronously. This is a major step backwards, and quite frankly it feels to me like trying to program in assembly. It's not expressive at all. You have to write your program in some pseudo code, then translate that to async code. And what for? What's the advantage you'll get? scalability? Who says you will need it? This is exactly where you should remember that premature optimization is the root of all evil.
To be fair, the exact problem is not async itself, but forcing CPS (continuation-passing style) for serial routines. For example, gevent and eventlet use greenlets (coroutines for Python) to avoid unnecessary callbacks in serial routines.
It might be a bit more performant than Django (my go to framework), but the amount of time it took me to produce the same work made it impractical.
The one web app situation I would use it for is when I have to make a request to a couple of databases/caches/etc, and each response isn't dependent on the others. That would make node far more performant than a sync framework. But this situation rarely seems to arise.
Basically, async I/O gives you more options than "block the whole world while you go read this stuff", and that means that old idioms aren't effective.
You gain more control: can choose when to block, when to limit concurrency and when to just launch a bunch of tasks at the same time. At the same time, you need to adopt a few new patterns, since you can't/don't feel right blocking execution every time you use an external data source. It's definitely a tradeoff and not a magic bullet.
For my longer take on this, see http://book.mixu.net/ch7.html
In fact, when you program for Node it's really important to keep this in mind, since (contrary to another statement from the article) not all libraries are asynchronous. If you select a synchronous db driver or write a long-running loop, it will block the rest of your program.
In general, though, I thought it was a good piece. I'm sure many heads have exploded on first introduction to node (and JS in general).
There is no such thing for node, unless you use a compiled addon. Node has no blocking networking facilities.
FWIW, Wikipedia seems to believe that "concurrency is a property of systems in which several computations are executing simultaneously, and potentially interacting with each other".
Regardless, sorry if sloppy (or poorly defined) terminology obscured my point.
"In computer science, concurrency is a property of systems in which several computations are executing simultaneously, and potentially interacting with each other. The computations may be executing on multiple cores in the same chip, preemptively time-shared threads on the same processor, or executed on physically separated processors."
A multi-threaded program running on a single core system has only one path executed at any given time, but we still call it a concurrent program.
"Two different code paths, can't do DRY" really? https://gist.github.com/1678395
"Oh noo, I can't return the results because they are async". That's what callbacks are for. You know what you CAN do? Do I/O in parallel that's what! Node makes it easy. https://gist.github.com/1678415
Anyway I hope this illustrates the point. The guy says it exactly right in one place: "Once you get your head around thinking in async terms, node.js starts to actually make a lot of sense." And therefore it is not a giant step backwards.
There are more elegant ways to write this (see http://qbix.com/plugins/Q/js/Q.js) but these are just minimal changes to his own code.
I would refactor the code to something like (still not as simple as synchronous):
asynchronousCache.get("id:3244", function(err, myThing) {
var useResult = function(err,_myThing){
// We now have a thing from DB, do something with result
// ...
}
if (myThing)
useResult(null, myThing);
else
asynchronousDB.query("SELECT * from something WHERE id = 3244",useResult);We'll start with the synchronous example:
myThing = synchronousCache.get("id:3244");
if (myThing == null) {
myThing = synchronousDB.query("SELECT * from something WHERE id = 3244");
}
This is verbose and tedious. We should really make the API look like: myThing = database.lookup({'id':3244}, {'cache':cache_object});
Let's apply this idea to his asynchronous example. We want the code to look like: database.lookup({'id':3244}, {'cache':cache_object}, function(myThing) {
// whatever
});
So instead of writing this: asynchronousCache.get("id:3244", function(err, myThing) {
if (myThing == null) {
asynchronousDB.query("SELECT * from something WHERE id = 3244", function(err, myThing) {
// We now have a thing from DB, do something with result
// ...
});
} else {
// We have a thing from cache, do something with result
// ...
}
});
We need to refactor this. Remember, node.js is a continuation-passing-style language. So let's set a convention and say that every function takes two continuations (success and error).Then, to compose two functions of one argument:
function f(x, result, error)
function g(x, result, error)
To: h = f o g
You write: function compose(f, g){
return function(x, result, error){
g(x, function(x_){ f(x_, result, error) }, error);
}
}
(Data flows right-to-left over composition, so "do x, then do y" is written: "do y" o "do x".)Now we can cleanly write a complex program from simple parts. We'll start by creating a result type:
result = { 'id': null, 'value': null, 'not_found': null }
Then, we'll implement cache functions that take keys (as results of this type) and return values (as results of this type). Looking up an entry in cache looks like: cache.lookup = function(key, result, error){
new_key = key.copy();
cache.raw_cache.lookup(key.id, function(value){
new_key.result = value;
new_key.not_found = false;
result(new_key)
},
function(error_type, error_msg){
if(error_type == ENOENT){
new_key.not_found = true;
result(new_key)
}
else {
error(error_type, error_msg);
}
});
};
Looking up an entry in the database looks about the same. The key feature is that the "return value" and the "input" are of the same type. That makes composing, in the case of "try various abstract storage layer lookups in a fixed order", very easy. (Yes, the example is contrived.) dbapi.lookup = function(key, result, error){ ... };
Now we can very easily implement the logic, "look up a value in the cache, if it's not there, look it up in the database": cached_lookup = compose(dbapi.lookup, cache.lookup);
cached_lookup(1234, do_next_step, handle_error);
You can, of course, generalize compose to something like: my_program = do([cache.lookup, dbapi.lookup, print_result]);
Writing clean and maintainable code in node.js is the same as writing it in any other language. You need to design your program correctly, and rewrite the parts that aren't designed correctly when you realize that your code is becoming messy.Continuation-passing style is pretty weird, but you do get some benefits over the alternatives. Writing a program with coroutines involves deferring to the scheduler coroutine every so often, littering your code with meaningless lines like "yield();". Using "real" threads is even worse; your code looks like single-threaded code, but different parts of your program are running concurrently. (Did you share any non-thread-safe data structures, like Java's date formatter? Hope not, because you won't know you did until the production code dies at 3am.) Continuation-passing style lets you "pretend" that you are executing multiple threads concurrently, but the structure of the code ensures that only one codepath is running at a time. This means that libraries that don't do IO don't have to be thread safe, since only one "thread" runs at a time.
All concurrency models involve trade-offs over other concurrency models. But when comparing them, make sure you're comparing the actual trade-offs, not your programming ability with each model.
I think you'll write better JavaScript if you know Python because Python encourages you to use named functions instead of lambdas. JavaScript fanbois get very excited about anonymous functions and overuse them; Python doesn't let you use anonymous functions for anything useful, so you tend to name things. (object.method is also nice syntax for working with callbacks.)
Anyway, Python and Node feel about the same to me, except for the fact that Python has nicer syntax.
It's certainly not a path I would relish having to follow, and I would not expect everyone to be able to program at this level.
I'd certainly not go so far as to criticize their programming ability for not being able to intuitively do this.
async.waterfall([
function(callback) {
async.map(ids, db.getById, callback)
},
function(posts, callback) {
callback(null, posts.map(templating.render))
}
], function(err, results) {
console.log(results);
});
EDIT: Fixed example. for blogPostId in recentBlogPostIds
await asynchronousDB.getBlogPostById blogPostId, defer(err, post)
templating.render post
Though you still can't simply return the result — you'd have to use a continuation — but it does make it simple enough to use CPS in general.This whole thing additional to the whole mess JavaScript is? I never liked it to begin with, but there is no real alternative until Dart is ready. Every time something comes out for JavaScript it adds another abstraction and chaos in my opinion. jQery for example, really impressive to begin with, but when you see what a horrible mess you can create with it...
There is a reason why big companies never adopt these things, i can't imagine how it would be to take over a node.js app from someone else.
Now you can argue that this takes practice. Crockford may write JS from heaven, but i don't want to invest my time in this language. These inconsistencies are not fun to deal with and when Dart is here, companies will drop it very fast.
I am now stuck with Scala, it is the complete opposite. It is complicated to get in, but when you get it, you have a gigantic toolbox to solve every problem the way you want. For web programming i recommend Lift, but when you want to get in fast and a fan of async try Play2.0. Node.js made async popular, it should get credit for that.
That's incredibly inaccurate. Big companies adopt a lot of crazy things, craziness isn't much of a deciding factor. Node.js is used by plenty of large companies and despite your tastes, Javascript in general is ridiculously popular in companies of all sizes.
asynchronousCache.get("id:3244", function doThing(err, myThing) {
if (myThing === null) {
asynchronousDB.query("SELECT * from something WHERE id = 3244", function(err, myThing) {
// We now have a thing from DB, do something with result
doThing(err, myThing);
});
return;
} else if(err !== null) {
// Handle error.
return;
}
// We have a thing.
});What a let down.
Someone needs to fix this
I advocate polyglot programming, and using node.js for tasks other than what its designed for (server side async programming) might result in unfavorable results.
iPods are not lousy because one can't text with them.
But now, he's not so sure.
Did I miss anything?
ETA: I understand that coding for Node looks and feels weird, but so does coding for Lisp, Smalltalk, Haskell and a long list of other programming languages.
Now, the thing Node has going for it, is that, despite being inferior than other similar technologies, it has a big following (community matters), lots of libs (libs matter), and it's based on an easy and familiar language to many.
You're exposing a synchronous API, but still can't take advantage of the huge ecosystem of Ruby libraries that already expose synchronous APIs.
Wrapping libraries becomes a one-off chore. Each individual library must be wrapped to work in an em-synchrony system, and if the libraries aren't both asynchronous and wrapped in fibers you can't use them. This not only shrinks the ecosystem of libraries further, but is also more error-prone than providing a general coroutine abstraction around socket IO.
Providing a generalized abstraction for doing synchronous I/O with sockets/fibers and an evented backend is exactly what I'm working on in Celluloid::IO:
There are libraries out there that add syntactic sugar to make async code look like sync code. Like,
group (
asyncfunc1()
asyncfunc2()
asyncfunc3()
) getPosts (ids, cb) ->
res = []
stash = (err, post) ->
res.push post
cb(res) if res.length is ids.length
db.getPostById(id, stash) for id in ids
getPosts [...], (posts) ->
# go on...
use `res[i] = post` if there is some implicit ordering. getFromCache = function (id, callback) {
asynchronousCache.get(['id', id].join(':'), function(err, myThing) {
if (myThing == null) {
asynchronousDB.query("SELECT * from something WHERE id = $id", {id:id}, function(err, myThing) {
callback(myThing);
});
}
else {
callback(myThing);
}
});
};
getFromCache(3222, function (myThing) {
console.log('myThing:', myThing);
}); getFromCache = function (id, query, callback) {
asynchronousCache.get(['id', id].join(':'), function(err, myThing) {
if (myThing == null) {
asynchronousDB.query(query, function(err, myThing) {
callback(myThing);
});
}
else {
callback(myThing);
}
});
};
getFromCacheSomething = function (id, callback) {
var query = buildQuery("SELECT * from something WHERE id = $id", {id:id});
getFromCache(id, query, callback);
}
getFromCacheSomething(3222, function (myThing) {
console.log('myThing:', myThing);
}); getFromCache = function (id, query, callback) {
asynchronousCache.get(['id', id].join(':'), function(err, myThing) {
if (myThing == null) {
asynchronousDB.query(query, function(err, myThing) {
callback(err, myThing);
});
}
else {
callback(err, myThing);
}
});
};JavaScript is single-threaded, so parallelism within a single JS VM is not possible.
What do you do if you want to see if, e.g. there are any results for the user's query?
I would do
if (results.length > 0)
or something