Why S3 went down
status.aws.amazon.com
status.aws.amazon.com
This is my favorite sentence of the day.
Now every time I see that damned whale on Twitter I'm going to find myself involuntarily crying out, "Oh, no! Failed while Gossiping!"
Everyone has outages: it's what you learn from them that counts.
If they know how it happened, it's not reflected in this article. The solution they describe addresses the detection of and recovery from future mysterious occurrences rather than identifying, understanding and eliminating whatever bug or condition caused this one.
> we're adding checksums to proactively detect corruption of system state messages
I dobut that means they're actually adding them together, they're adding checksums to the process to ensure data corruption has no effect. Or am I misunderstanding you?
Trying this approach with consumers could backfire horribly if someone misunderstands the details or tries to twist them out of context.
That's a lot hard to do with a developer audience.
By 2:20pm PDT, we'd restored internal communication..."
I wonder why it took 3 hours to clear state.
It didn't -- it took 3 hours to get all the machines talking again after the state was cleared.
This is probably mostly due to two factors: 1. When doing a "clean restart" Amazon almost certainly had each S3 node look at the data it had stored to make sure that they had correct metadata; and 2. After each node had been restarted, Amazon probably had to relink them gradually rather than all at once -- most self-organizing network protocols have limits on the rate at which nodes can join or leave the network.
This question no verb? Or rather... I have no clue what you're asking, can you clarify?
For example:
Jack: I wonder what Jill's favourite flavour of ice cream is.
Sam: Chocolate?
Jack: Yes, that's the one.(sorry, I had to...)