Fixing the callback spaghetti in node.js
github.com
github.com
I've been using LuaJIT embedded in Nginx (LuaNginxModule). Lua supports coroutines, so a function can just yield. Here's a brief example:
-- Query a database using an http backend.
-- Yields and handles other requests until the reply is complete.
local record = util.getUserRecord(userId)
-- Send some text to the client. Yields control while
-- the actual transfer is in progress.
ngx.print( "Result=" )
-- Send the result, encoded as JSON, to the client
-- Again, this call doesn't block the server.
ngx.print( cjson.encode(result) )
With code like the above I can easily handle into the thousands of concurrent connections per second on the lowest end Linode VPS node available, with barely any load on the box -- and I'm told it should be able to handle 40k+ connections per second, if I were to do any tuning. Oh, and I have only 512Mb of RAM, which it doesn't even get close to under load. And the longest request took less than 500ms at high load.I've been using OpenResty [1] which has the Lua module and a bunch of others all configured together. Works great, and I can't complain about the performance.
Someday I'm sure I'll hand the maintenance of this off, and then I might regret not using one of the "popular" frameworks. But the code is SO straightforward using this stack -- and what I'm using it for is so simple -- I think not.
I had several failed private projects doing the same on C, Lua, and Haskell. Haskell is pretty ideal but I’m not smart enough to hack the GHC. Lua is less ideal but a lovely language - I did not like that it already had a lot of libraries written with blocking code. No matter what I did, someone was going to load an existing blocking Lua library in and ruin it. C has similar problems as Lua and is not as widely accessible. There is room for a Node.js-like libc at some point – I would like to work on that some day.
V8 came out around the same time I was working on these things and I had a rather sudden epiphany that JavaScript was actually the perfect language for what I wanted: single threaded, without preconceived notions of “server-side” I/O, without existing libraries.
http://bostinnovation.com/2011/01/31/node-js-interview-4-que...
I don't think node.js is good enough because you have to deal with the issue being discussed here. In Haskell, as you said, you just write normal code with no worries about mutable state messing with your thread.
https://developer.mozilla.org/en/JavaScript/Guide/Iterators_...
It's the same situation when using callbacks.
I don't have a lot of experience in this area. I'm just reporting back what I remember hearing.
That is true for a single threaded reactor, but in practice I haven't found it to be an especially useful property because naturally that doesn't include external state (i.e. the database).
Getting atomicity right in evented-code has in fact often been a hairier issue for me than doing the same with co-routines or threads, because when you're not allowed to block ever then you quickly find yourself in a situation where you need a retry-mechanism.
Such codebases then tend to quickly converge towards the actor-pattern (tied together by queues) which, ironically, could be had much easier by starting out with co-routines in first place.
So except for things you should EXPECT to change (like the state of a database you're querying) between calls, the "world" you (as a programmer) should care about stays perfectly consistent.
Unless I'm not understanding what you're asking -- is there some situation that I haven't encountered where the state of something that isn't a local variable matters?
Of course people like to point at the overhead of native threads and assume coroutines have similar overhead, which is total bunk. Ironically, event sources in node use a stack-like tracking element which brings back a similar sort of overhead you see in coroutines in Lua, for example.
I doubt we'll see node take another shot at coroutines but that's okay. Node will do fine without coroutines but it will come at the cost of making certain types of code a little less natural (same as eventide code in threads becomes unnatural).
Wondering, is the OpenResty solution otherwise single-threaded-with-an-eventloop much like node?
From the nginx wiki:
> Unlike Apache's mod_lua and Lighttpd's mod_magnet, Lua code written atop this module can be 100% non-blocking on network traffic as long as you use the ngx.location.capture or ngx.location.capture_multi interfaces to let the nginx core do all your requests to mysql, postgresql, memcached, redis, upstream http web services, and etc etc etc (see HttpDrizzleModule, ngx_postgres, HttpMemcModule, HttpRedis2Module and HttpProxyModule modules for details).
Is that talk-via-nginx-commands thing cumbersome in practice?
local doc = ngx.location.capture( "/couchdb/hamster/"..uid );
This queries the local CouchDB for a particular user record, and stores it in a local variable.My mental model of code was one-dimensional: do this, then do that, then do that. So I was used to exception handling for errors, and my programs were like trains on a track. This was a comfortable abstraction, but the way in which my programs were written did not reflect what they were doing in the real-world: reading from disk, reading from the network, waiting for something, doing something, responding to something, receiving something in chunks.
Now, looking back, I prefer writing code that reflects what my programs are doing. There is more headspace, another dimension, no longer one thing after another, but a stack of things hovering and happening at the same time, interacting with each other, moving forwards through time.
Now, instead of trying to make concurrent work appear non-concurrent, I prefer to embrace concurrency and see how I can write for concurrency. This is almost certainly different to the way in which I would have written synchronous code. Code for me is less complex now and shorter now than when it was synchronous. It feels richer and more descriptive of the work being performed. It's also faster and more reliable. In essence, I have learned to write better concurrent code.
The first example from this page shows a request handler initialising a database connection and then executing a query. That's terrible separation!
Callbacks "spaghetti" actually does a great job of highlighting when you're not abstracting enough, any more than about 5 indents and you should be seriously considering refactoring your approach.
Thus
app.get('/price', function(req, res) {
db.openConnection('host', 12345, function(err, conn) {
conn.query('select * from products where id=?', [req.param('product')], function(err, results) {
conn.close();
res.send(results[0]);
});
});
});
becomes something like app.get('/price', function(req, res) {
products.fetchOne(function(err, product) {
res.render('product', { product: product });
});
});
Also, as other posts on this page mentioned, try{}catch{} is not how errors are handled in Node.JS, plenty of async operations will gain a fresh stack, and cannot be caught in this way.You comfortably omitted the error handling (admittedly the original snippet didn't have that either).
However in production code it's not optional, and that's where node.js code tends to get really nasty.
app.get('/price', function(req, res) {
products.fetchOne(function(err, product) {
if (err) {
res.render('error', { error: err });
} else {
res.render('product', { product: product });
}
});
});There you nailed the problem. It's a constant headache, especially since library authors have different ideas about the format and semantics of "something useful".
Before you know it you have re-invented your own half-baked exception-framework, to normalize/wrap/re-throw those ErrBacks. And then a week later you notice that rollback/retry and event cascades need a whole new level of treatment.
There's a reason why mature event-frameworks such as Twisted have never quite taken over the mainstream. It's sad to see node burn all its powder on a niche programming-style that just doesn't fly for the majority of applications.
> Callback spaghetti is a sign that you're doing something wrong.
I have some Node.js code that does direct uploading via Amazon's S3 multipart uploading API - * multipart form processing, callbacks
* each part requires multiple S3 API calls, callbacks
* parse XML results from the API, callbacks
Granted not all workflows are this complex. But many are - and they will result in callback hell. But saying that people are doing something "wrong" is at odds with the reality that complex workflows are a fact of life.I'm not saying you don't need to have a series of nested callbacks to do things, I'm saying you should hide these behind the appropriate abstractions for the task you're writing.
In your case the bullet points listed are exactly those layers of abstraction.
The request handler processes the form and calls the S3 layer for each file. The S3 layer then calls the APIs and passes off the responses to the XML parsing layer - which gives it back a useful JS object detailing the response, the API layer then provides a response to the request handler which formats the response and sends it to the client.
The workflow I'm implementing in Node is even more complicated than this, but good use of abstractions and control flow libraries[1] means that it's extremely rare for any code to be indented more than 5 blocks.
[1] https://github.com/caolan/async is my lib of choice, but plenty of other suitable choices exist.
This is not going to catch most errors occurring in an asynchronous APIs:
try {
asyncOperation(function(err, result) {
// ...
});
} catch (e) {
// ...
}
Most errors will occur asynchronously, thus the convention of "err" being the first argument to asynchronous API callbacks in Node.Unless this supports that convention, your code should actually look like:
async function magic() {
try {
// code here
await err, bar = doSomething();
if (err) {
throw err;
}
// more code here
await err, boo = doAnotherThing();
if (err) {
throw err;
}
// do even more stuff here
}
catch (e) {
// handle the error
}
}
...or something similar. It's better than the alternative, but not great.This certainly could support the "err" first convention, but APIs that don't use that convention wouldn't work correctly.
// these two lines are equivalent
await a, b, c, d = moo(1, 2, 3);
moo(1, 2, 3, function(a, b, c, d) {} );
So I don't see why not... // ...these two lines would be also equivalent
await error, foo = bar(1, 2, 3);
bar(1, 2, 3, function(error, foo) {} );Here's a concrete example. Normally in Node you do something like this:
fs.readFile('/etc/passwd', function (err, data) {
if (err) {
// handle error
}
// normal processing
});
The await version would look like this: await err, data = fs.readFile('/etc/passwd');
if (err) {
// handle error
}
// normal processing
Ideally it would look like this: try {
await data = fs.readFile('/etc/passwd');
// normal processing
} catch (error) {
// handle error
} try await foo = bar(baz)
as shorthand for await error, foo = bar(baz)
if (error) throw errorUsing only that, I rarely go over 80 character column limit that I impose on myself. There is absolutely zero callback spaghetti whether it's 2 or 25 functions deep in the chain.
Tbh, callback spaghetti only happens to newer async programmers in the same way that a newer programmer will write arrow code with if/else statements.
It's simply not a problem that needs to be addressed other than educating people who are new to node.js with some example tutorials that use an async helper library.
Now, let's say c() changes and needs to call someAsyncFunction() and provide a callback. Which means c itself needs to take a callback. Which means b needs to provide a callback, so b needs to also take a callback, so a needs to provide a callback. And so forth.
Callbacks are infectious - once anybody in the call stack needs one, everybody needs one even if they just pass it on down the stack. Unless you don't need to do anything with return values, but that's fairly rare.
It looks pretty decent in CoffeeScript:
blah = (done) ->
async.series [
(next) -> foo next
(next) => @bar.baz 1, 2, 3, next
(next) ->
x = y
z w next
], doneAt least that's my experience with structuring asynchronous event-driven programs (without coroutines).
Without it small build/utility scripts turn into macaroni cheese unless you resort back to Python/Ruby/Bash - more languages, harder to maintain.
http://tomasp.net/blog/csharp-fsharp-async-intro.aspx
http://news.ycombinator.com/item?id=2999260
http://bvanderveen.com/a/owin-buffering-async/
http://devtalk.net/csharp/async-await-and-c-vnext/
http://msmvps.com/blogs/jon_skeet/archive/2010/10/29/initial...
(read comment)
Anyway, it shouldn't be too difficult to build the compiler yourself and set up a basic emacs/vim + mono environment for linux.
koush just took this to the next level.
But you know what. Just stop fighting the callback model. Adapt your coding idioms and move on...
"Programs can be automatically transformed from direct style to CPS." [1]
Do the math.
function foo () {.... }
function bar () {....foo(); }
function barfoo () {....bar(); }
doSomething(barfoo);this.io.sockets.on('connection', this.onConnection.bind(this));
(disclaimer: I am the guy who worked on the now-defunct `defer` support in Coffeescript, and am now working with the onilabs folk on the stratifiedJS runtime)
For example, F#'s let! binding works with any monad not just async.
Using await - is the program flow suspended until the corresponding function returns , or does await keyword act more like a 'pause and continue' mechanism.
You will pry Python from my dead, cold hands :)
I've written a LOT of async code in Perl and C, and ran across this problem countless times. It really is a blessing to not have to worry about it.
The attraction of using Javascript server-side for many is the possibility of using the same language client- and server- side, with the relative ease of a native data encapsulation (JSON) for transferring state between the two. Even back in the VB6 days (I say "back in the days" ruefully: we are still supporting the product!) I used JS for a large chunk of validation code server-side, this meant that the same logic was used in both places aside from some wrapper code to arrange for the code on both sides to be presented with the same data model. From the point of view of a relative beginner to server-side coding who is adept at Javascript on the client side this is particularly attractive (much like Python being available in browsers would be to seasoned server-side Python programmers)
The attraction of node in particular is a mixed bag:
* Momentum. People are using it so people are using it.
* The event driven nature, if properly handled, can make it very memory efficient.
* The speed of V8, which at the time was a step or two ahead of other javascript engines (since node first turned up there has been plenty of active development in this area, with a different story each month about one or the other JS engine beating everything else in some benchmark or other, so I'm not sure which has any sort of upper hand at the moment).
* It arrived at the right time, and development from "proof of concept" to a relatively mature product happened fast enough that people didn't become disillusioned soon after the initial excitement.
* Decent library support, partly due to the number of people created the above mentioned momentum. It is easy for an alternative server-side stack to suffer in this area and eventually die because of it. Though the rapid development may make this bite back a bit, as a lot of bindings from six months ago that have not been actively maintained might not work because of changes in node over that time.
* Other Javascript based server-side stacks have started to die off already, such as Jaxer (http://en.wikipedia.org/wiki/Jaxer#Aptana_Jaxer) for one example, partly due to the popularity of node.js though in many cases it was already happening (or they didn't gain much traction to start with).
http://tamejs.org/ also has a good write-up.
If you're going to innovate, then design a language that compiles down to JS that provides the innovation. CoffeeScript.
If you want to take that a step further, look at ClojureScript. Want delimited continuations? Fine. All w/o requiring you to fork Node.js or CoffeeScript.
foo.step1 = function(){ do_some_thing(foo.step2) }
foo.step2 = function(arg){ do_another_thing( ... ) }
foo.step1()
# enjoyRemember that not everyone is as awesomely brilliant as you are.
Writing a user interface that can respond to user input, whilst also being able to handle and respond to a long running data access or computational requests is a concurrent problem.
Having a single threaded model with callbacks, like the JavaScript browser model is one of the less complex ways to handle this.
I agree that a less complex model is appropriate for some developers/applications. But to call yourself a "skilled-professional GUI-developer", you need to get awesomely brilliant enough to handle this.
We'd like to think that the C# guys were looking our way when they came up with async/await, but there's no proof. :3
====
/* https://gist.github.com/1250314 */
var db = require('somedatabaseprovider'),
compose = require('functools').compose;
compose.async(getApp, connect, select)({ url:'/price', host:'host', pass:'123' }, function(error, shift){ shift.conn.close();
shift.res.send(shift.products[0]);
});====
It's that easy to abstract those messy callbacks using some functional tools.