Using Flat Files So Elections Don’t Break Your Server
open.blogs.nytimes.com
open.blogs.nytimes.com
I'm curious about their provisions for cross-datacenter failover. The article mentions haproxy being ready to direct requests to a different datacenter as well as ELB spanning availability zones. I'd expect a failover option entirely outside AWS as well, with short-TTL DNS ready to make the switch.
I'm also not sure what value varnish brings to the table when the Apaches are just serving a small number of flat files. It seems like unnecessary complexity -- by the same logic found earlier in the article -- when a well-tuned webserver will serve flat files at a comparable rate. Maybe the Apache configuration for this workload was sufficiently different to make it an unwanted risk.
In their "flat-file" setup though - they are deliberately configured so that, should varnish fail (which was one of their concerns) - they don't actually need it - they could just have haproxy immediateley start hitting the apache servers directly.
They say they ran it on a EC2 micro instance, and the micro instances are specifically designed to handle spikes, not continuous heavy load. In fact, they intentionally throttle under continuous heavy load.
I guess it must have been the right choice for them, I was just surprised to see it was a micro. Would that typically be the right way to go?
Apache can, one way or the other, do most of what the others can do - possibly not as well, and at the risk of a much more complex configuration.
Haproxy does one thing and does it really,really well - it's great for dealing with load balancing, concurrency limiting per user defined resource, and identifying and routing incoming requests to the right infrastructure. It's really good at this - that's what it does.
Varnish does one thing and does it really, really well - it caches content and serves it up (usually out of memory, but even if it's swapped, it's optimized) to keep the load off your application servers. It has various optimizations built into make it really good at this.
So - in this case, the answer might be "not much - we could just put the apache's out front behind haproxy" - but it appears putting the varnish server out front with a 5 second cache dropped the load on the application servers (in this case apache serving static files that it receives over rsync - dont' forget the rsync part - resources are needed for that). This might result in smoother output for the end user, rather than something hitting a node that's busy servicing an rsync update. It may also be their engineers are very familiar with the haproxy/varnish front end setup, as it presumably exists in their current day to day operation as well.. so the people responsible for keeping things up probably decided "Yes, we'd like to keep it there - it makes our lives easier."
There is an operational anti-pattern in there - removing too many elements from a known system is also a kind of added cmoplexity - all your troubleshooting methods disappear.
Their goal here was to de-couple the dynamic elements from the event-driven side of things and turn them into something more resilient (and less flexible) for a short time to deal with unknown and unpredicably large load.
A flat file is essentially a .csv holding data, and can be a fast way to bulk load denormalized data. As such, it actually does have a use in the context of scaling, so it's natural to expect that they were using the term correctly.
A static file, or more simple, a HTML file is what they're actually talking about. As they've noticed, it's what web servers are best at serving, and it scales obnoxiously well.
Now that we're all talking about the same thing, I can say that I've been serving all my product blogs as .html for years, and have never had any of them fall down under load.
It's really easy to get something in place to generate static files from a blog or CMS. I do it with:
- a 404 handler that maps missed requests for .html to their equivalent generator.
- a regular old blog engine that takes an extra parameter "writeThisToHTMLOnceYouveRenderedIt"
Future visitors skip the redirecting and generating and are simply served the static file. Next time you edit the content, you can simply blow away the whatever-blog-entry.html and index.html and know that they'll show up again next time anybody asks for one of them.
Static yes, HTML no, an HTML file could still have e.g. SSI instructions. It's valid HTML, but if the webserver supports SSI it's not going to be static.
"All the Code That's Fit to printf()"
(I'm also a big fan of Varnish. I tried it out this weekend on a site that basically serves an HTML file that says "hello world". Apache can do about 11,000 requests a second, but Varnish can serve 15,000 requests a second. Excellent!)
I did not use varnish at the time, but I basically kept up with the load. My pages were cached with an in-memory cache in the app layer, which worked well enough, I guess. As blog.jrock.us mentions, I was unhappy with the design of my software, so I took it down. Two years later, I almost know what a good design is, and so I should have a blog again soon. But I digress :)
If you're using linux using IPV4 have a look at the following parameters in /proc/sys/net/ipv4:
tcp_tw_recycle
tcp_tw_reuse
Tuning those will help in allowing faster re-use of sockets in the TIME_WAIT state. This matters because at the defaults your sockets will linger for a long time before being allowed to be re-used (as per the RFC). Technically this is the correct behaviour but it can quickly become a bottle-neck.
Ulimit max open files per process:
ulimit -n 50000
The default is just 1,024, and that's not nearly enough to keep varnish working hard.
Those are the first things to look at, there are many more once you start to bottom out again, but this will get you started.
A good way to see if you've got your kernel tuned properly is when varnish starts to approach the limits of your hardware, I've seen it do well over 600Mbps on an otherwise unloaded box.
How about cache it once at write-time, then again whenever you edit it?
>After each batch of new data was received from the AP, this server determined which pages needed to be re-rendered and, using the Typhoeus libcurl-multi bindings for Ruby, pulled new data for each of these pages from the render pool.
Sounds like their infrastructure already has a farm of servers set up to handle rendering articles into their HTML components, so for this scenario they would save the output as a static file and then write it to disk rather than sending the HTML out to the response / caching layer / etc.
"Flattened files" may have been what they were meaning. "Flattened Files" would be a nice simple term for them if there wasn't confusion with Flat-file databases.
to be honest, since there wasn't actually a problem to solve as varnish is setup as both an HA environment and to use the grace/saint features, this is a case of overengineering in my book.
Static content is easy to crank up to web-scale. You can use DNS, any kind of load balancer, all kinds of web services, CDNs, whatever.
Dynamic content is hard to scale (compared to static). You have several layers of added complexity. Yes, you can build wonderful, self-scaleable systems - but at some point they hit a limit, there are many more resources that can be tied up, and troubleshooting and scaling that out beyond anything you've previously imagined on short notice can take time you can't afford.
So - simply de-coupling the dynamic content generation from the static web serving is a great way to make a clean break - you now have a known & tuneable load on your dynamic application (because your'e running it at known intervals, rather than being event driven by user requests) and you have a front-end static infrastructure that you can scale like mad, and even if your back-end collapses, edit by hand.
Surely there are other ways to approach the problem.... but it also depends on the engineers involved, the time taken, and their confidence in their ability to deal with it.
I'm also fairly sure they aren't the first company out there to take this approach to burst scalability issues... but it's curious to note how the NYT actually operates.
TL;DR: Look at the old configuration, and the new configuration. Decide which one will best serve your business in terms of your ability to troubleshoot it when it gets hit by a level of traffic higher than you can plan for, because you have NO idea how high it will go.
As the author points out, having the Times election site go down on election eve would be a BIG PROBLEM -- massive losses in both reputation and advertising revenue. For the system architect, a failure could possibly mean losing his job.
What the author has laid out is a system that is robust to multiple, simultaneous failures (with possible exception of the loss of AWS, although that's not entirely clear). That just seems like good planning.
From the man page:
"sendfile() copies data between one file descriptor and another. Because this copying is done within the kernel, sendfile() is more efficient than the combination of read(2) and write(2), which would require transferring data to and from user space."
For the regular ebb and flow of data, you can plan, and use varnish/rails/nginx/ all the magic tools you want, and things will work great - but you also have added complexity, and when something goes wrong, the more complex the system, the longer it generally takes to fix. Especially under unexpected heavy load.
Simplifying the system down to flat files and de-coupling the ROR stuff gives you a clear troubleshooting point - if something goes berzerk, you can cut the link between the two and troubleshoot in relative safety before turning replication back on.
I'm rambling - the main point is to reduce the complexity involved in the end-user transaction to be as efficient and fast as possible so you can deal with an unknown load factor coming in on a really important day. Going down that day would be BAD for business.
Generate the static files: those files contained live election results, so they had to be regenerated as new results came in.