What if Microsoft and Danger #Cloudfail turns out to be a #SANfail?
siliconangle.net
siliconangle.net
As far as I can tell, "cloud" generally means someone else stores my data and the only way I can get to it is via the Net. In that case it's absolutely a failure of the "cloud".
And if the data was stored on a SAN, it's a SAN failure. If someone didn't do backups at all, it's a backup failure.
What's with all the labeling? Look, if I pay someone else to manage my data, and they screw it up and I suffer data loss, it doesn't really matter what piece of gear or level of process is responsible -- someone still screwed up.
GFS/HDFS are officially called filesystems, but they're about as much as a filesystem as a Samba server: they expose files to the outside world, with some extra layers of abstraction you can mount the filesystems under Linux, but in essence they live on top of existing filesystems themselves and have nothing in common with the traditional filesystems, where you operate on a block device. A SAN, otoh, is such a block device and traditional filesystems operate on them.
This means that you can run a RDBMS on a SAN, while you can't on GFS/HDFS. This covers a common use case where you have a central expensive database server, with dozens of terrabytes of storage in a SAN, which you see in a lot of large corporations.
Backup works like any other system; you do it app layer. In high-volume high-sensitivity apps, the people running the disk server also back it up, by running protocols that mirror block-level changes to a second rack full of disks somewhere in Tennessee.
(yes, yes, or IDE or FC, yes yes, or software initiator, etc etc).
Sorry but "cloud" or not, SAN or not, the MS/Danger fail was a stunning failure of process. Who cares what the underlying technology was or how it failed? What matters is that a multi-billion dollar software company actually managed to loose, permanently, their customers' data through obvious carelessness, through either lack of a backup in some fashion or other (and don't even try to argue they could have had an excuse - "multi-billion dollar company" "reputation" backups might be sort-of hard maybe - perhaps - but MS is supposed know what it's doing). Lightning and Asteroids wouldn't strike five different carefully chosen locations...
My guess is that this will hit MS really hard over time. Even if they actually were hoping Danger would dry-up and blow away, they've now done the worst case scenario to customers. Repeat after me "never let MS near your data...".
I was in complete agreement with you up to this point. Microsoft's strengths have long ago shifted towards marketing. They once had a edge in technology - over Lotus, over Borland, over Novell, over IBM and even over Apple (in the OS7~9 days).
Sadly, after their market share - and dominance of the OEM channels - exceeded a certain tipping point, they had really no need to excel in technology. From that point on, what differentiated Microsoft was a careful management of the expectations of their core market (VARs, OEMs and CTOs).
At Microsoft, technical competency has taken a back-seat since the early 90s.
This entire mess up smells like "Management by numbers". When a manager cuts costs without considering the implications to data-safety, this is what we get in the end.
Seriously: how much data needed to be backed up? 50 terabytes? 500 terabytes? And why the hell the data resides in a SAN and not distributed and replicated in a bunch of front-end servers? The SAN should be the back-up.
I love the idea of having all my data on the cloud - It's a major nuisance to switch cellphones and port data from one to the next - and having the data on the cloud could make it easy to access it through a nice web interface not unlike Google's GMail, Calendar, Picasa or Documents. Implementing "back-up your data" and "restore your data" links somewhere would also be quite easy.
And yes. Never let MS near your data.
All along critics have been saying that you'd be nuts to trust someone else to back up and secure your important data.
You trust people with the right, next generation trust worthy architecture.
When you go for leading edge you often end up with a mix of unproven stuff, assorted approaches, uneven performance and problems nobody ever had. It's fine to experiment, but you should only trust your customers' data to new technology when you find a combination that works well all the time (or, at the very least, when you have mapped all the circumstances it doesn't).
This whole disaster could have been avoided if there were no single point of failure - at the very least, three identical clusters being brought off-line and rebuilt from mirrored data for the upgrade one at a time in order not to disrupt service to users. I am astonished a telco did not know this.
No matter WHAT the deployment mode was, the lack of backups is what has turned this into an issue, NOT the SAN or "Cloud" failure.
We will all be renting computing and storage in the short future.