Despite the size of 22 double-XL instances, they were a small portion of our overall footprint in EC2; it takes well over that kind of capacity to run the platform.
Those instance types all happened to be in one availability zone for a variety of reasons. Our platform overall does not live in a single zone.
Losing machines is not a problem for us (we cycle them constantly, in fact). Normally losing even that many machines would not even be noticed by our customers; this was an unusual case in which several factors cascaded into a larger problem.
To be clear, this downtime (45 minutes or so, with full normal state by 90 minutes) was, unfortunately, our fault - not Amazon's. EC2 instances vaporizing is an expected part of using the service.
We've made a couple of operational changes that will prevent these issues in the future, and we sincerely apologize to any customers who were affected.
That being said, given how many toy, no-traffic apps are likely running (for free!) on the platform, I think we can safely assume there's a very high degree of multi-tenancy.
I have no firm data, but a customer last year put a very low volume web app on Heroku, and I guess this is exactly what we saw: the first request after a quiesant period took many seconds - then subsequent requests were quick. If I am right about this, customers with high volume web apps would not notice this.
34.2 GB of memory
13 EC2 Compute Units (4 virtual cores with 3.25 EC2 Compute Units each)
850 GB of instance storage
64-bit platform
I/O Performance: High
Which means each application has 16MB of ram. Does chroot (or whatever jail mechanism they use) allow shared memory for standard libraries?