Of course once you get beyond the headline, I think most people are much worse with protecting themselves from rare outages than Google.
Of course once you get beyond the headline, I think most people are much worse with protecting themselves from rare outages than Google.
Normally, Google redundantly distributes out data to at least 3 different geographically distinct locations. Check out the 'BigTable' white paper [0] for more info.
For 99% of cases (and pretty well all user cases), this would not cause data lose. The key here is that the data was generated on the servers and did not have a chance to duplicate before the event.
link here: http://static.googleusercontent.com/external_content/untrust...
actually in colossus one can tune RS coding parameters per file, to get a tradeoff between performance/durablity.
RS coding uses less copies, but same level of safety (tradeoff is the recovery computation time.)
EDIT: In this video https://vimeo.com/100153741, around the 23 minute mark
Unless I am specifically paying for backup service, I wouldn't expect/want them to do that for me. And even if I was, I'd still have offsite backups if the system was important.
If you aren't being responsible with your backups in a noisy environment like Google Cloud/AWS, understand that you are vulnerable to freak accidents like this. Google/AWS's job in all of this is to try to reduce the frequency of issues and to minimize the impact.
Google Compute Engine offers customers the option to make snapshots for backup, or use a true "cloud" storage engine. If anyone lost data here it was customers explicitly not doing backups and only using a single zone. I don't know why anyone would expect different. GCE easily allows you to network machines in multiple data centers, but close geographically. So you'd only need to handle region-wide disasters.