Mr. Moore gets to punt on sharding
37signals.com
37signals.com
However, the inherent design of 37signals applications make sharding a lot easier. With 37signals apps, you're not dealing with all the users in the system or all of the data. You're dealing with one company's data and the users from that company. You could much more easily put data from companies with names A-I in DB1 and J-Z in DB2 and call it a day than most sites can shard. You wouldn't have to worry about things like losing cross database joins or having to lookup in multiple databases for certain reads.
In fact, with Basecamp, there's 6 possible domains that your site can be a subdomain of (projectpath.com, clientsecton.com. . .). Each one of those could point to a different set of app/DB servers. If you ran metrics on usage, you'd have to run it 6 times and aggregate the data, but in terms of user usage, each could be its own little separate system not even knowing the others existed.
Sharding becomes a problem when you can't fit all the data that will be used together on a single server. For example, Facbeook could shard by college, but you're allowed to have friends at different colleges from your own. So, you have to do things like cross shard lookups and that mess. With 37signals products (to my knowledge), everyone gets an account with your company's site and interacts with your company's stuff and if they want to interact with another company's stuff they log out and log back in with a different username/pass. So, everything can be neatly sharded by company since you don't have things that cross that boundary.
Eventually, you get into some decently ugly stuff if you get big enough. Most people never have to worry about sharding (even 37signals doesn't), but there aren't just easy answers.
Using in memory tables, flat files, etc where possible beats databases.
Your previous setup: Webhead1-|- Database1 Webhead2/
So, you have two webheads reading from one database. You need to shard and have a second database: Webhead1--Database1 (Companies A-M) Webhead2--Database2 (Companies N-Z)
Easy! Think of it this way: because a person at Company A is never accessing Company Z's data, there doesn't need to be any communication between the databases. So, it shards in the same way that Wordpress shards: different sites are on different servers.
With Basecamp, you don't need every company to be on the same server because you don't move data between companies. But you don't need a separate server for every company either. So, you can put a bunch of companies on the same server and a bunch on another server and you're done. There's nothing magic or RoR specific.
Or monkey-patch activerecord yourself to use a shard-index.
I'd only run one application server too, if I could get away with it. Dealing with 1 is much easier than dealing with many. Do it as long as you can!
I'm not going to say "this guy does not know what he is talking about", but what kind of shitty database are you running if you need to have the entire database cached in memory?
Isn't the point of having a proper RDBMS that it should be able to handle data more efficiently than what would be "logically" possible?
I've run databases with 100s of gigs of content on a 2-node cluster where each node "only" had 8GBs of RAM and performance nor memory usage wasn't an issue at all.
One point of real pain we’ve suffered, though, is that migrating a database schema in MySQL on a huge table takes forever and a day.
Oh right. MySQL. That explains it.
His IO-problem is because he is using a shitty database. Only in the MySQL-world is it common to need the same amount of RAM as the dataset you are working on.
And if you are going that way memory-wise, you might as well you use an in-memory flat textfile or persisted objects.
I'm not saying having excess RAM for caching is bad, because that is obviously silly, but when you need to have the entire DB in ram, you should realize something is horribly wrong.