> A lot of it was EBS and also the replication system, which turned out to be breaking due to a bug in our own code when the whole thing slowed down due to EBS.
This bug was a pretty rare case, but to be specific, here it is:
1. Databases for a particular data type are specified in the config file like:
dbmaster, dbslave1, dbslave2, dbslave3
2. In Python this ends up just like that: a list where the 0th element is special. There's a "config list" (the list read out of the config file) and an "active list", which is those from the config file that have recently reported themselves as live.
3. Reads are sent to (sort of) random.choice(active_list) and writes are send to active_list[0]
4. But wait! If the master is "down" (the EBS backing him has slowed down so far as to mark him unresponsive) then active_list[0] may very well be a slave
5. Londiste now can't reconcile the slave database with the master database. There's a sequence conflict!
The symptom would be that a replica would refuse to replicate, we'd detect the replication lag on that slave, and remove it from the active list. Then its load would fall, we'd get an alert, and we'd say "well crap" and rereplicate to it.
In practise this didn't happen very often, I can only think of a few occasions off-hand, and generally the master's disk lag was so bad that a downed slave was the least of our problems.