> In practise this didn't happen very often, I can only think of a few occasions off-hand, and generally the master's disk lag was so bad that a downed slave was the least of our problems.
It got progressively worse as time went on the EBS performance got worse. It was happening every couple of weeks after you left, and I think it got even worse after that.
Luckily we had dug into the londiste internals and managed to figure out how to restore replication much faster, but that was still super painful. :)
But you're right, the disk lag usually caused the immediate problem -- the out of sync slave was usually pretty easy to deal with, depending on how many got out of sync at once. It was just annoying and time consuming.