Why Node.js streams are awesome
blog.dump.ly
blog.dump.ly
There's at least one more downside: the user loses all indication of progress as the Content-Length is unknown when the headers are sent
dumply knows the exact size of each image as it is saved in the DB on upload, and all the zip byte headers are fixed, so the zip file size should be deterministic and calculable even before the first byte is sent. Remember we don't compress the already compressed images.
If you didn't know the file sizes, for example you had raw unknown input streams, or had compressible data, you can still guesstimate the content-length so the user got some progress bar, even if it wasn't 100% accurate.
Looks like using browser sniffing, you can deliver an exaggerated Content-Length to everyone but Opera and browsers will deal with it gracefully. Pretty neat. (Obviously not desirable for violating the HTTP spec, but the UX gains might be worth it.)
For example major mobile operators pipe HTTP connections through a proxy that recompresses images. In that case you see e.g. Safari's or Opera's UA string, but you're actually dealing with proxy's HTTP behavior.
An older article (2008) that also talks about misreporting content-length for fun and profit: http://tech.hickorywind.org/articles/2008/05/23/content-leng...
Are suggesting they would rather guess the browser by inference through presence of DOM properties and methods?
In the end it narrows down to: would you rather a) break the functionality for some users in exchange to give the best possible solution to others, or b) give an OK experience to everyone
I usually go with "b". I think that frustration is much more powerful than awe.
But yeah, it's probably indeterminate behavior.
server = http.createServer(function(request, response) {
request.pipe(ffmpeg.process.stdin);
ffmpeg.process.stdout.pipe(response);
});
What a nice interface! The end result is that you can do weird stuff like: $ curl -T my_video.mp4 http://localhost:9599 | mplayerhttps://github.com/evanmiller/mod_zip
The module is quite mature at this point, and is used in production on many websites (including Box.net, which commissioned the initial work). The module supports the Content-Length header, Range and If-Range requests, ZIP-64 for large archives, and filename transcoding with iconv. Being written in C, it will probably use much less RAM than an equivalent Node.js module.
I have found that the hardest part of generating ZIP files on the fly has nothing to do with network programming; it's producing files that open correctly on all platforms, including Mac OS X's busted-ass BOMArchiveHelper.app.
Having a large number of stream primitives means you can easily wire up endpoints, for example say you wanted to output a large db query as xml, or consume and editing gigabytes of json, or consume, transcode and output a video.
You can by all means write a nginx module in C for each usecase and this is probably the right solution for very HEAVY specific loads.
But writing a C module is probably a barrier too high for many, whereas implementing a nodejs stream isn't. Respond to a few events, emit a few events and you have a module that can work with the hundreds of other stream abstractions available. (npm search stream)
You still need the specific domain knowledge (eg how zip headers work) and this is usually the complicated bit. mod_zip looks excellent, and I wonder if some of the domain knowledge of handling zips can be resused in zipstream.
This is how we handle it currently.
> User adds images to a virtual lightbox.
> User decides that he wants to download all the images in this lightbox, so presses "Download Folder". The user is then presented with a list of possible dimensions that they can request.
> The user selects "Large" and "Small" and hits "Download"
> This request gets added to our Gearman job queue.
> The job gets handled and all the files are downloaded from Amazon S3 to a temporary location on the locale file server.
> A Zip object is then created and each file is added to the Zip file.
> Once complete, the file is then uploaded back to Amazon S3 in a custom "archives" bucket.
> Before this batch job finishes, I fire off a message to Socket.io / Pusher which sends the URL back to the client who has been waiting patiently for X minutes while his job has been processing.
This works okay for us because when users create "Archives" of their ligtboxes, generally they do this because they want to share the files with other people. This means that they attach the URL to emails to provide to other people.
So for us, it's actually neccessary to save the file back to S3... however, I'm sure that not everyone needs to share the file... it would definitely be worth investigating if the user plans to return back to the archive, in which case implementing streams could potentially save us on storage and complexity.
With streams, there is no need to cache, as recreating the download is dirt cheap. Essentially just a few extra header bytes to pad the zip container, ontop of the image content bytes that you will have to always send.
The use case you mentioned, of sharing the download link, works exactly the same. You send the link, and the what ever user clicks on the links gets an instant download.
True you are bufferring data through your app, instead of letting S3 take care of it. But if your on AWS, S3 to EC2 is free and fast (200mb/s+), and then bandwidth out of EC2 costs the same as S3. If it goes over an elastic IP, then a cent more per GB. You app servers also handle some load, but nodejs (or any other evented framework) live to multiplex IO, with only a few objects worth of overhead per connection.
In return, you can delete a whole load of cache and job control code. Less code to write, test and maintain.
The cost when streaming and not streaming should be pretty much the same, unless your non-streaming case is working on-disk (in which case you're comparing extremely different things and the comparison is anything but fair)
http://wiki.nginx.org/X-accel and mod_zip for instance.
Why do people keep reinventing the wheel, thinking node is the be all and end all when this is nothing new at all?
It should be possible to do something similar using e.g. generators (in Python) or lazy enumerators (in Ruby)
In fact, in Python's WSGI handlers return an arbitrary iterable which will be consumed, so that pattern is natively supported (string iterators and generators together, then return the complete pipe which will perform the actual processing as WSGI serializes and sends the response). Ruby would require an adapter to a Rack response of some sort as I don't think you can reply an enumerable OOTB.
Using eventlet or gevent was much kinder to the system.
It's just in node, doing it in the evented way was actually simpler and quicker to implement, than the 'ghetto' way. This isn't usually the case, and I always recommend doing the simplest thing that works first. It's just nice here that the simplest thing is also a tight solution.
Our original solution was literally a shell exec, and it was perfectly fine (..for a while)
I feel like from a UX perspective, it'd be ideal to be able to give some friendly error message to at least acknowledge that the failure is on the server end. A page that says, "Sorry, we're having trouble accessing your files right now. Please try again in a minute.," seems more user-friendly to me than a download that suddenly fails with no explanation. Nevertheless, this is very cool.
[1] http://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html section 14.35.
The main benefit here isn't that it's possible to do this thing, as many people pointed out the myriad ways this is accomplished elsewhere. The key point is that everything that manipulates data, node core as well as the userland libraries, implement the same interface.
The notion of piping the output from some I/O (say, a request to S3) into the input of some other I/O (say, a currently writing HTML response) without ever referencing it is blessed by node, which has a stream type as part of its standard library. But it's far from the most common abstraction of asynchronous work.
public static void myEndpoint() {
HttpResponse resp = WS.url("url-of-file").get();
InputStream is = resp.getStream();
renderBinary(is);
}
Or am I missing something?(EDIT: This doesn't do the zipping or multiple files -- I guess I need a ZipOutputStream to take it the rest of the way)
No - the image requests happen on the server, which then concatenates them together in to a zip file which is served to the client. The client never sees the actual image files, just the resulting zip file.
I'd imagine (well, hope anyway) that the ISP proxies that downscale image files do so based on the HTTP Content-Type header - since the images contained within a zip file would be part of a file with a different Content-Type they should be left alone.
Many other languages have similar abstractions — python and generators for instance http://www.dabeaz.com/generators/ — although their usage would likely require more work as they probably are not as standardized as far as usage goes.
I mean in this case it's "faster" to write it because somebody else had already gone through the motion of creating a zipping stream (which they still needed to fork), it's not like node magically did it.
tldr: the node community is re-discovering dataflow, and a few are trying to pass it as some sort of magical property of node.
Theres nothing specific about node (architectures scale not frameworks), I'm just saying that node makes this kinda of architecture relatively simple. Use the correct tool for the job.
also checkout lazy haskell lists, which provide a very powerful (and related) abstraction at the language level.