How Ravelry Scales to 10 Million Requests Using Rails
highscalability.com
highscalability.com
Tokyo Cabinet/Tyrant is used instead of memcached in some places for caching larger objects. Specifically markdown text that has been converted to HTML.
and this one tip:
The database is the problem. Nearly all of the scaling/tuning/performance related work is database related. For example, MySQL schema changes on large tables are painful if you don’t want any downtime. One of the arguments for schemaless databases.
Not much "how" in that.
http://news.ycombinator.com/item?id=802889 (http://www.tbray.org/ongoing/When/200x/2009/09/02/Ravelry)
Looking over the setup, that's quite a lot for one person to do during nights and weekends in 4 months. Pretty cool.
Edit: yes, Casey says in the Tim Bray interview that "As soon as we could, we got alpha testers in to try it out... 4 months later, we had a site that we were ready to announce."
10 million server requests per day sounds kind of impressive, until you actually do the math.. divided by how much physical iron they're using, that's a little less than 9 requests per second per server.
It makes me wonder: if they were using something other than rails, would they need that much iron?
Rant: I wish technical sites would stop using req/day as a metric. It leads to the op type of analysis. At the very least, such articles could use a format of "X req/day peaking at Y/s". Maybe if the NYT was writing it would be ok to use req/day but a sight who's tagline is: "High Scalability Building bigger, faster, more reliable websites." should know better.
(You can't save the sixth server if you want your site to be up while the seventh one is rebooting or being replaced.)
If the other comments are to be believed, this site was built by one person, working part-time, in four months. He can't afford to lavish time on unimportant problems, like desperately trying to conserve server resources that he could otherwise afford and that cost far less than a programmer's time is worth.
Of course, you're also totally throwing away the "new agile" platforms, like django, scala lift, and so forth.
2.5 of our physical servers run Passenger/Rails.
If this were a Java app, I definitely would have been able to get away with less (mostly because of less memory consumption, but less CPU consumption wouldn't hurt either)
However, I'd probably still want 2 machines for redundancy.
http://programming-gone-awry.blogspot.com/2009/06/how-to-sav...
They could use Nginx -> HAProxy -> Nginx(w/ passenger), but the Apache version feels slightly more mature (e.g. it has some config options that the Nginx version lacks) and it's likely they were already using it before the Nginx version came out.
http://www.igvita.com/2008/12/02/zero-downtime-restarts-with...
It's really a great piece of software. Kudos to Willy.
PS - you're also correct about the nginx->haproxy->apache. nginx makes a fabulous front end and I just plugged in Apache/Passenger where Mongrel used to be. I like that 1) I can easily plug in something else in the future and 2) Passenger on Apache is very stable. Nginx support is newish and I'm running stripped down Apaches that only do Passenger, so I'm not too fussed about it.