How I want to write Node: Stream all the things
caolanmcmahon.com
caolanmcmahon.com
Streams, honestly, are hard to keep straight when the program gets big without a stronger type system. IMHO.
Some really sharp people have been working on stream computing software in Haskell for a while - Gabriel's Pipes package is a good example of generalized stream computing with strong equational reasoning as its foundation.
Maybe if you really want to try and do this in Node you can gain some inspiration from his journey: http://hackage.haskell.org/package/pipes
cat some.txt | sort | uniq > yay.txt
Then it isn't a problem - it's obvious and simple, but I think there will be difficulties as the programs get larger and type-level awareness occupies more space in the programmer's brain vs. it being handled by the compiler...Node follows the unix way. Everything is a stream. It's just Buffers and JS objects flying around. It's really stupid, and sometimes it's nasty. This isn't helped by Javascript's warts.
But there's an enormous upside to this: following the stupid Unix way means that no matter what you need to do with your data, there's an npm module for it. Just .pipe() your stream in and your code is done. This is amazing. And it's possible only because of how bare-bones and loose the Buffer stream API is.
Strong typing has its place, but it would ruin node's biggest selling point. It's hard to realize this without trying it.
I understand the benefits and that's why Pipes in HS is such an exciting thing because it gives us a formally reasoned and general set of stream computing tools - you can compute anything with type-level guarantees. It's just as flexible and general as, say, Unix pipes but better because there are guarantees of the library's tooling and there are guarantees of the programs you produce! You can't say that in Node / JS, Python, Ruby, etc...
JS's lack of strong typing limits your ability to reason about streams (a lot more than just streams, too) and further limits your ability to write performant stream computing software. Pipes, in Haskell, give you the big three: Effects (I/O), Composition (function composition with fusion), and Streams (generators and iteratees); because of the type system Haskell (and some nudges here and there by the library author) can fuse and optimize that code to a ridiculous degree in addition to all of the other nice guarantees you get from the type system (separate of I/O from pure code, etc...)
I personally don't think dynamic typing is a selling point, ever - I write software faster and with fewer bugs in Haskell than I ever have before in Python, Ruby, Erlang, or Scheme. But that's a totally different topic and I don't want to derail this one.
Don't misunderstand me as being aggressive, please. I fully respect what people decide to like and work on, I'm just trying to expand the awareness that there are tools in existence that do it better.
I think you missed my point so I'll restate: in practice you don't reason about streams in Node, because the community (a product of the simplicity of the streams API) has a packaged solution to your problem. It plugs right in. And this ecosystem exists because of the simplicity and dynamicity of the constructs used.
I actually agree with you that Haskell does it "better". It's purer and cleaner. You'll probably have less bugs if you write everything in Haskell.
Except it doesn't matter to me, because Haskell doesn't have anything close to the plug-and-playability of npm modules -- and this is a pure social product of the stupid interface that Node exposes compared to Haskell. Node is shittier, and that's why it's more capable at solving the problem I have -- constructing powerful apps in close to no time, and zero lines of my own code.
I guess what I'm saying is that sometimes worse is better.
I would be pleasantly surprised if this existed at all in the Haskell community. Even more so with two lines of my own code.
But the API provided is trivially replicated, and, yes, easily built atop pipes.
The absolute most compelling thing about Node is how it has hacked organizational dynamics in large companies. Walmart and PayPal basically used it to completely liberate their frontend groups from their backend systems using facade system with huge improvements on customer systems.
What is sad is that all the other high concurrency systems are going to end up implementing much of the GHC runtime without the reliability of Haskell...
I know Caolan's been thinking about this and reworking it for a while, so I'll be interested to see whether it manages to see significant takeup.
- promises: about 20% of the room.
- async: about 80% of the room
For me async.waterfall([list of functions]) is little nicer than 10 chained .thens().
And people advocating promises still keep saying it keeps things flat. No it doesn't, we're already flat because we're all using async. Stop pretending async doesn't exist and isn't massively popular.
And way, way better documented. Q.spawn what? And this is the best promises library?
Stack Overflow question: Simplest fs.readFile example with generators and Q?
Current only answer:
Q.spawn(function* () {
…
var data = yield Q.ninvoke(fs, "readFile", somefile);
…
});
Answer from Highland docs: var data = _.wrapCallback(fs.readFile)('myfile');
- What's Q (yes it's a module, but what does it mean? Is it supposed to a misspelt queue or something else?- What does 'ninvoking' something do?
- Shouldn't I just be able to to put the variable declaration outside of the scope?
- Why do competing Open Source implementations of the same standard exist? Can't there just one reference implementation?
That's not the future.
I might be really ignorant here. I probably am - I could read a shit tonne of docs to work out what this strange beast does and technically someone can probably do a better job answering that Stack Overflow question. But nobody has, because very few people know how to operate the current state of the art generators/promises setup.
From the Q docs: "If you have a number of promise-producing functions that need to be run sequentially"
No, I don't have a number of promise producing functions. Nobody in nodeland has that. I just have functions. I could read about turning them into promise producing functions, and calculate whether this abstraction layer is adding value, but then again, I could do productive work with async.
And from the looks of it, Highland too.
async.waterfall([
fn1, fn2, fn3
], function(err, result) {
});
Q: resultP = [ fn1, fn2, fn3 ].reduce(Q.when, void 0);
resultP
.then(function(result) {});
.catch(function(err) {});
If you chain thens in Q, you are not doing it right (imho).The biggest value to me is being able to avoid the passing around of callbacks and them relying on varying conventions (some async are function(args, ..., callback(err, data)), some are function(args, ..., callback(data), errback(err)), some are function({success: callback, error: errback})).
Promises solve this problem by not passing around callbacks _at all_. Instead you return the promise and let the consumer attach the callback itself. And once we have promises widely available and part of the standard library, the calling conventions will be standardized too.
EDIT: I agree that as it stands, the lack of standardization of promises (jQuery's are mutable, for instance) is a pain, and that documentation could certainly be better.
var Promise = require('bluebird');
var fs = Promise.promisifyAll(require('fs'));
Promise.spawn(function* () {
var data = yield fs.readFileAsync(somefile);
});
Documentation: https://github.com/petkaantonov/bluebird/blob/master/API.mdMy favorite example is doing a diff using an async diff service which provides a function `svc.diff(string1, string2)`. But also imagine that you need to preprocesses the files using the sync function `removeEmptyLines(buffer)`. This is how the function looks like when implemented using Bluebird:
function diffTwoFiles(f1, f2) {
var file1 = fs.readFileAsync(f1).then(removeEmptyLines),
file2 = fs.readFileAsync(f2).then(removeEmptyLines);
return Promise.join(file1, file2).spread(svc.diff);
}
I'd love to see someone come up with a better example using callbacks and async.- Isn't data a scope down?
- https://github.com/petkaantonov/bluebird/blob/master/API.md is API based, not task based. async is API based too, but the API has names like 'waterfall' and 'parallel'. I can click them because I know what they mean.
Promises just makes me feel like I'm reading about and endless series of abstractions.
I didn't quite understand the comment about data being a scope down. What do you mean?
Yes, promises do have quite a steep learning curve :/ However they're a lot more flexible than a utility grab-bag of functions that never quite fit the problem you're having. By that I mean I often had to massage my functions (by creating new closures or using bind etc) to make them fit the signature that async requires.
I wrote a bunch of examples for common tasks here - http://promise-nuggets.github.io/
For me, nearly everything .waterfall(), .each(), or .parallel(), and has been for two years now.
Promises are not about syntax sugar. They're about utilizing the whole power of the language and providing a parallel for most features found in synchronous code:
1. Functions have return values
When using node style callbacks, we're ignoring the fact that the language was designed with functions that have return values. Instead we use half-functions. Its no wonder those compose quite badly - the language wasn't designed for that kind of composition. The language was designed to work with functions that take input values and return an output value. Callback-based functions do only the first half. Thats why to get them to compose we resort to a bunch of hairy helpers and boilerplate code.
Callback-based functions that don't return anything are seriously crippled in power, and promises fix that, restoring much of the power.
2. Errors can bubble like exceptions
When using node style callbacks, we must explicitly handle all errors. On one hand, this is a good thing: we should deal with all errors. On the other hand, its quite tedious: most of the time we can't deal with the error at the exact place it appears but must pass it up one level in the call chain.
Promises do the error bubbling automatically. We can attach the appropriate error handler at the appropriate place to deal with the error.
This simple feature results with tons of useful patterns, one of which is the ability to manage resources with constructs like C#'s `using` keyword. - https://github.com/spion/promise-using
3. Values in variables can be accessed multiple times
When using node style callback and event emitters, we must make sure to "capture" the value exactly when it comes. If we don't do that, poof, its gone forever - we missed it.
In contrast, promises will keep the value for us. If we need to access that value later, we can simply attach another callback handler. An example where this may be useful is a database connection:
We initialize the connection and get a promise for that connection:
var pConn = db.connect(host, port);
How do we implement a query method that is immediately available and will queue up queries until the connection is established? Easily: function query(q, params) {
return pConn.then(function(conn) {
return conn.queryAsync(q, params);
});
}
It doesn't matter whether the connection was established a long time ago or hasn't been established yet - the query will either execute immediately or its execution will be delayed until the connection becomes available.Now try doing this with callbacks :)
Promises can be used as flow control, but more importantly, it's an object that encapsulates asynchronous mechanics.
I like to see how async can launch an asynchronous operation, then allow listeners to be attached later to capture the result. Now, you may say that if you want to attach listeners to capture result, you'll want to use event emitters. True, but event emitters has its own problem because event emitting and attach listeners are synchronous. What if the event was emitted before any body has a chance to attach listener to it?
Promises does not have these issues.
Where the line checks for `window.Promise`
var outcome = Boxon();
fs.readFile( "test.txt", outcome );
// Attach listener later:
outcome( function( err, content ){ ... } );
See https://github.com/JeanHuguesRobert/l8/wiki/AboutBoxonsBoxon objects are light compared to promises, but they interop well. Best of both worlds!
In short, its API is nearly as extensive as Q's with essentially no overhead—bluebird is hardly more expensive than callbacks, while Q is something like 10x slower than using callbacks.
I ask because I wrote some Lisp macros (I work in a Lisp that compiles to JS) to implement a few async patterns I need, and making sure that exceptions are trapped and threaded into the callback chain correctly was the most complicated part.
From a practical perspective, it doesn't make sense to try to catch exceptions in asynchronous code, anyway. Once you do something asynchronous, you lose the stack and thus the try block. The way to catch thrown exceptions in asynchronous code is with domains, which something as low level as async would not be expected to handle.
Sure, but what if the error is thrown at you as an exception in the first place—which happens a fair amount, because that's how the JS runtime tells you when something is wrong? How do you get from there to the callback way?
What the Lisp macro I mentioned does is generate a separate try-catch around each block of code that runs at a different time and thus might throw an exception that would not otherwise get caught. In that way it catches every exception that's thrown, converts it to an error object, and passes the error back through the callback chain. The async library could do the same, albeit with a lot more code. I'm curious why it doesn't.
From a practical perspective, it doesn't make sense to try to catch exceptions in asynchronous code, anyway.
I don't think that's right. Asynchronous code is just synchronous code that runs at different times. Each block of synchronous code can generate exceptions. I agree that if you don't catch them then, they become useless; but you can catch them then. The reason this is not a "practical perspective" in JS is not that it doesn't make sense, it's that the language doesn't support it. Even the minimum code necessary to catch every exception involves so many try-catch blocks as to obscure the rest of the program. So no one writes such code by hand in JS.
Yet it is, I think, code that one wants, because without it you don't have a consistent error model. You end up having one model for first-class errors—the ones you detect and pass to callbacks before an exception has a chance to arise—and a second one for the dregs—the ones that come from any code that didn't know about or follow the callback convention (which, critically, includes the language runtime). The latter kind of error either crashes the server or gets caught by a top-level handler so it "only" crashes the request it was processing. That's a half-baked system.
Its on you to catch that, not your libraries. This shouldn't be terribly common, though. The only thing I can remember having to wrap in a try/catch in the codebase I work on is JSON.parse.
> The async library could do the same, albeit with a lot more code. I'm curious why it doesn't.
It couldn't, without domains. try-catch wouldn't do it. Domains are something that is not very well understood, in my experience, and expected to happen at a higher level than libraries like async.
As for domains, I don't know what you mean by them, but if they're catching errors at a higher level than the async library, my guess is that they must be some more sophisticated sort of top-level handler; perhaps something that keeps track of which async calls are in progress and attempts to bind exceptions back to their context? Whatever it is, it sounds complicated.
But what I understand least of all is how you guys all seem to write Javascript code that generates almost no exceptions. To me that sounds almost like bug-free code. No null references, for example? I get stuff like that all the time.
We write in CoffeeScript, where a null reference check is so astonishingly easy to write that you use them everywhere you might get a null. I'm not sure what other exceptions you're seeing. We do basically no math, so /0 errors aren't a problem.
> For example, I don't know why you say that the async library couldn't try-catch every place that an exception might occur
Let's build a typical function you might pass to async:
function(next) { request.get(url, function(err, data) { JSON.parse(data); } }
Let's assume the server doesn't serve JSON like we expect - so JSON.parse throws an exception. The only thing async could have wrapped in a try/catch is the main function, but we've fired off a request and then the call stack wrapped up, including the try/catch. Next, an event occurs that calls our callbacks, not going through async at all. That's where the exception occurs. The stack trace generated by that exception doesn't contain any code in the async lib, so it can't possibly have a try/catch active.
Domains are a way of fixing this. You create a Domain and bind callbacks to it - if that callback throws an exception, the Domain instead emits an error event.
Re null reference checks, to get behavior analogous to a null exception you have not only to check for null, but also pass back an explicit error if you find it. That's a lot more work than adding in an extra question mark. Null checks that do nothing but not crash are a mixed blessing; 90+% of the time they do what you want, but when they don't, you get a silent failure and a debugging goose chase. I'd be surprised if you told me that that never happens.
I took a look at Node.js domains and they do seem really complicated. If I were working in Javascript instead of having control over the language, I doubt I would use them; I would probably just crash-and-restart as one of the other commenters described. That's not a good solution, but probably the best tradeoff given the alternatives.
That said, we get very thorough testing from our large user base, and we quickly fix crashers. Our server proc crash rate is almost 0, brought up by occasional spikes on releases.
Boxon.cast( fs.readfile, 'myfile' )
.then( function( content ){ console.log( content ) })
.catch( function( err ){ console.log( "Error", err ) });
Boxon is promise implementation agnostic, works with Q, bluebird, etc... See https://github.com/JeanHuguesRobert/l8/wiki/AboutBoxonsWith these improvements to streams in Highland making things more convenient and broadly applicable I expect to be using Highland streams for certain things.
Sometimes simplicity is a feature, too, though.
Is there some way to know what kind of source you've got, or are the the sources constructed in a way that chooses which behavior you get?
Rx doesn't handle automatic back-pressure (like Node Streams) but does have mechanisms to avoid overwhelming slow consumers. Rx also has delayed subscription which you can call lazy, but not by turning the stream into a pull-stream (allowing you to sequence actions in the way Highland does).
If any of the above needs further qualification or comment please weigh in on the issue by commenting here... but for now I'll leave it at that. I actually list RxJS in the blogpost because it's a good example!
Still fleshing it out, but pretty close to calling it complete: https://github.com/Reactive-Extensions/RxJS/tree/master/src/...
We're more than open to pull requests though if anyone thinks we're missing something here.
- We have object.defineProperty() in ES5 to avoid enumeration.
- You can use user-specified prefixing to avoid future conflicts.
Eg:
{foo: 1, bar: 2}.hlPairs();
rather than: _.pairs({foo: 1, bar: 2});But you can call it whatever you want, so... who cares? The choice is yours... Being a pedant is hardly constructive.
I'm excited about this, as someone who uses underscore.js and async together, heavily.
"A lot of the functions" (versus "all of the functions") sounds like a one-way ticket to readability hell, since it means that an inattentive reader may assume the functions ARE from Underscore.
When a new thing comes out, the discussion should focus on what's significant about it.
You're describing the state of affairs before Underscore existed, back when functional-ish programming in JavaScript was ruled by Prototype.js:
... which added a lot of useful methods to native prototypes.
While handy in controlled and limited environments, mucking about with native prototypes quickly becomes extremely dangerous and difficult — once you have two third-party modules on the page that expect different versions of your patched prototype method ... once you have a new version of a browser that implements one of your previously-extended functions, but does it differently — you're pretty well screwed. Both of those things tended to happen in large sites.
You can have ten different versions of Underscore loaded on the page, living in peace and harmony, in ten different third-party modules. Not that you should. But that you could.
But to be perfectly fair, you are 100% correct: this is not a technical complaint and perhaps this entire sub-thread is, as has been claimed, "bike-shedding." Since this lib is designed to be loaded via an AMD-style mechanism, users can call it whatever they want, so _ is just as valid an identifier as any other. Except the obvious issue that readers of the code, examples, and any code that follows suit, will end up with this completely pointless ambiguity because _ is ultimately a meaningless name if it evolves to mean simply "some library I loaded." You may as well write sample code like:
var $ = require('http'); $.createServer(...);
I would think people would call that out as ridiculous and confusing.
While Underscore is in a better position than those that extend native prototypes, regarding api/environment conflicts, it can still be tripped up by poor shims because it defers to many ES5 methods if they exist. For example, if Prototype 1.6.0 and Underscore.js are included on a page Underscore's `_.reduce` method won't work properly. This is one of the reasons why libs/frameworks like Dojo, Ember, Lo-Dash, RequireJS, Sizzle, and YUI do native checks too.
Which is not to say that it's a good choice, but instead to say that the badness stems from overloading (+) too much and having free coercion.
(In one of the ways of constructing the natural numbers within ZFC set theory in mathematics, one identifies 0 with the empty set {}, 1 with {{}}={0}, 2 with { {}, {{}}} = {0,1}, and so on.)
Saner perhaps, but still ungodly complex with weird edge cases.
LazyJS just received stream support, but I'm pretty much sold on to highlandjs.
Why is this so an enlightenment for node guys? UNIX does it right since epoch - simple programs perform simple tasks and connected via pipes. Python has gevent, so you don't even need 'a stream' or other bullshit, you just write the code as-is and greenlets provide the concurrency needed.
The real enlightenment comes from 'programming properly'; you start with C and torture your brain with function pointers and realize why it is a good idea to treat functions as a first-class objects. then you learn some 'proper' functional programming languages like lisp or something to learn how to think in functional way. which is the only guaranteed and proven path to prevent yourself from shooting your own foot by writing 20+ nested callbacks. If you start with binding an anonymous function to a <button>'s click event and think you can do this to do real programming, you'll never get it right.