BigPipe: Pipelining web pages for high performance
facebook.com
facebook.com
It'll be interesting because this will require significant rework to fit with how most web servers work. It would be hard to implement in NGinX, for instance. Facebook probably just wrote a custom server, or heavily modified the code to an existing one.
This does kind of beg the question, why use nginx at all? It provides you with a lot of protection against malformed requests and general fuckery. If you need it you can push the middle layer back into nginx as a module which would be screaming fast.
This reminds me of Heroku who do something like this with their 'routing mesh.' client -> nginx -> routing mesh (erlang) -> thin (ruby app server). The erlang process knows which ec2 instance has the ruby process to serve the request and initiates and is basically a smart proxy. You could quite easily query 6 backends simultaneously for each pagelet and pipe the JSON out to the client then throw in the footer as well.
I think this deserves some experimentation.
nginx + erlang + rails (serving JSON).
Guys from Taobao (http://www.taobao.com) have open-sourced pretty much everything you need to do that:
- ngx_echo (http://github.com/agentzh/echo-nginx-module) for asynchronous pipelining,
- ngx_drizzle (http://github.com/chaoslawful/drizzle-nginx-module) for fetching data from Drizzle/MySQL/SQLite,
- ngx_postgres (http://labs.frickle.com/nginx_ngx_postgres/) for fetching data from PostgreSQL,
- ngx_rds_json (http://github.com/agentzh/rds-json-nginx-module) for converting database-responses into JSON.
Also, this isn't new concept, at least few Chinese companies I know of use similar rendering process.
I've prepared simple proof-of-concept configuration for nginx:
http://labs.frickle.com/misc/nginx_bigpipe.conf
As you can notice, every "sub-page" is generated individually. Using presented configuration everything is chunked and flushed, so it will be sent to the client right away. Response on the client side looks like this:
http://labs.frickle.com/misc/nginx_bigpipe.output
DISCLAIMER: I don't know how Taobao is using released modules internally or if they use them in production already (but I know some portals do).
The Javascript half of BigPipe catches those flushes, handles the dependencies, registers handlers, and slaps the HTML into place. This lays the groundwork for many things like incremental page updates, parallel execution, and so on.
So it may be a good candidate for a library.
I guessed they were doing something similar a few months ago when Facebook first started showing up block by block. Interesting to see it all laid out.
ESI got alot of attention about 10 years ago, but then sort of fell out of favor in the tech media/blogs. Some big companies are still using them like Akamai, and the Varnish HTTP accelerator has some basic support for it.
I always liked the idea of breaking up my page into smaller segments, and then caching each part independently, and assembling the page from the cache. The cache could request only the parts of the page that aren't in cache/expired/uncachable, but otherwise pull everything from a super-fast cache.
ESI is really cool for caching inside your infrastructure, but it doesn't help the client as much because they have to download the entire page again even if only the time changed in the top bar. I've always dreamed of a way of cutting up my HTML page to have different pieces 'cached' by the browser. A logical next step from this style (which helps you load a client with a cold cache) is to have the client cache each of these pagelets. On subsequent requests you could return JS (and actually you could check a cookie and recycle the JS too!) look in the client's HTML5 storage or some other crafty mechanism which will reduce strain on FB's infrastructure and make it faster for the user.
Client side includes didn't make it into HTML5. Maybe in HTML6? Check this space in 2020. http://lists.whatwg.org/pipermail/whatwg-whatwg.org/2008-Aug...
NoScript does not publish their individual add-on usage statistics, but the global download/usage ratio can be calculated from the statistics on the Firefox Add-on home page [3]:
1,962,617,946 add-ons downloaded 157,090,095 add-ons in use
About 8% of downloaded add-ons are still in use. Assuming NoScript's usage ratio is comparable to the average, approximately 5.4 million installations of FireFox are running NoScript. Let's ignore the fact that NoScript's usage ratio is probably much lower than the average, due to the fact that it breaks most web pages.
I'd wager that the average NoScript user has at least two machines, so the total number of NoScript users is probably less than 2.7 million.
There are over 230 million internet users in the USA [4]
Even if 100% of NoScript users were Americans, they form 1% or less of the general population. If you, like me, believe that I have been generous to NoScript here, it is likely that no script users are no more numerous than 1 in 1,000.
Even if the user is using NoScript, they can whitelist your site. You can probably add <noscript>WARNING: THIS SITE IS BUSTED WITHOUT JS</noscript> to the top of your page and call it a day. If you are feeling generous, redirect no-script users to the mobile version of your site and tell them why.
tldr: Assume that human user agents have Javascript.
[2] https://addons.mozilla.org/en-US/firefox/addon/722/
[3] https://addons.mozilla.org/en-US/firefox/
[4] http://www.google.com/publicdata?ds=wb-wdi&met=it_net_us...
EDIT: Converted to blog post -- http://news.ycombinator.com/item?id=1406233
Keeping two different output formats for a site (one for crawlers and one for humans) sounds complex. To date, none of the sites I've developed could have justified such overhead in development.
demo at: www.mixhammer.com
but using pages in chunks instead of only the assets on the page
At least, as described. Rendering most of the page immediately and then leaving the connection open to shove more through is more server-push than the standard ways of doing this.
Of course there's a lot more server-side and client-side work going on to support facebook's implementation here. They've essentially invented some crazy way of packaging multiple requests together but hiding that inside a broken html document.
If server push is the destination they I might have thought that websockets was a cleaner way of achieving it. I don't think server-push is the destination though in this case.