"Why not just plan ahead? Because most of the time, it was a very abrupt failure that we couldn’t detect with monitoring."
"Why not just plan ahead? Because most of the time, it was a very abrupt failure that we couldn’t detect with monitoring."
So you have a system, and you have monitoring in place. Let's say the monitors were set up for 1 minute polls, because somebody thought that was a good idea. Suddenly you find out one of your servers is down. Oh noes! There's 45 seconds until the monitor finds this out, which would be horrible.
Since we have doubled the reads on the existing servers, we now no longer have capacity and connections are stacking up. Shit :'( But not to worry! Let's just quickly kill the extra reads - now we have more capacity! Hooray!
Except, if the extra reads weren't happening, they would have already had extra fucking capacity and not had to flip a switch in the first place.
Now you see why i'm mad, bro?