Amazon CTO: 'You should be able to walk away' from cloud providers
networkworld.com
networkworld.com
This is a ridiculous statement. PUTs are $0.01 for 1,000 requests, and GETs are $0.01 for 10,000. Deletion is free. Data transfer out per GB is comparable to the monthly cost per GB.
Sure, putting data in is cheap because you don't get charged for bandwidth, but storing it costs just as much per month as moving it out would. If you've got somewhere better to store your data than S3, you'll come out ahead price-wise after a single month, so you should do it. That's not lock-in, not even close.
Convenience-wise, yes, there's a lot of lock-in with AWS, no question. But that's a different problem, it's not fair to imply that you're locked in because of charges to actually get your data out.
I think you're understating the cost of keeping such a backup, and the complexity that may come with it.
Having a backup: $1000 a year for a dedicated server you physically own.
Not all data can be backed up for only $1000 a year either; it's not just a matter of storage costs and a server. A write-heavy service like Foursquare can't just turn off and do a dump at night; to have backups they need enough server capacity to replicate data as it comes in. For those servers to keep up, they have to be as beefy as the ones they're replicating -- a year or so ago, that was 64GB RAM per database for their Mongo instances. That's definitely more than $1000 a year's worth.
You are mixing data, memory and processing power. In the case Amazon crashes and burns, they need that CPU+Ram+bandwidth etc. near the data.
But to just keep a copy of the data (not available online) they don't need a really beefy server -- just store all the updates from a node, and replay them later at your leisure. 5400rpm will often be fast enough for this purpose (of course, you would need days to recover in this setup ... but that might be ok if past data is not needed online)
Take twitter or foursquare for example: They need everyone's most recent max(4, tweets/checkins in the last day) available online at any given time. But if last years' tweets/checkins are not available online when the system is degraded in case of backup restoration etc -- that may be acceptable.
But even if your data does matter, spending $1000 to back it up separately is not a no-brainer, depending what your data is actually worth to your business and how likely you think it is that S3 will lose it. Assuming your data is worth a million dollars, then spending $1000 to back it up implies that you think there's at least a 1 in 1000 chance that S3 will lose it all.
I don't know what probability I'd assign to "S3 loses all my data", but I do know that the particular data I store there is not valuable enough to make it worth backing up elsewhere. If my data actually mattered, I'd work out the math and consider alternatives. I will say this, as a word of warning - I would never use Amazon's durability claims ("Amazon S3 is designed to provide 99.999999999% durability of objects over a given year") when doing such a risk vs. cost assessment, because the real risk is not that the expected failure modes will "get you", but that something totally unknown and unpredictable will happen, i.e. someone at Amazon totally fucks up a staging push and sends it live to production instead, and the replication system goes nuts and starts corrupting everyone's data irretrievably before anyone realizes what's going on, or something like that. SLAs and durability claims don't account for stuff like that, but it's still a risk.
Of course, there's also the "cover my ass" and "I'm not paying for this anyways" factors to consider - if you would likely lose your job if the data in question was lost, and your personal finances are not affected one bit by the company spending an extra $1k/$10k/$100k on backups, then even if math tells you it's a waste for the company, you might as well do it, because as long as it's not too much of a pain to set up it will help protect you. The calculus always changes when it's other people's money you're playing with...
But the real cost is the time people spend working on it. Sys admins, programming the process to get it onto the server, etc.
I don't think any of the data I have in s3 would cause my company to go out of business if we lost it. Most of them are thumbnails from videos that could be regenerated.
Btw if anyone has any suggestions I'd love to hear them. So far the best approaches seem to be either to upload to multiple providers in the first place or to organise the buckets in such a way that bucket listings are not required to copy all files changed since a certain date.
That's how I took it anyhow.
http://assets.en.oreilly.com/1/event/31/The%20Cloud_s%20Hidd...
http://www.slideshare.net/sh1mmer/the-clouds-hidden-lockin-n...
There were four separate things needed to bring about what we think of as the modern international mail system. First you needed better infrastructure inside and between countries. You needed literally more portable objects, envelopes with stamps on them, instead of loose sheets, scrolls, wax seals, etc. You also needed standardization of the rates and address formats so one letter can travel anywhere. Last was a uniform rate, no-questions-asked promise to deliver via optimized routes, what we would now call a "peering agreement". Two ends of the agreement would treat each other as peers, and honor each others' comminucations as they would their own.
It's that kind of system we don't have, but should, in the cloud. We need better infrastructure in the form of optimized routes between clouds. We need to be able to move our virtual machines and configurations around without special help. We need to make sure we don't get locked-in by screwball APIs or data formats. Most of all, we need the various cloud and web services vendors to commit to honoring each other's traffic without clobbering us, their customers, with metered billing.
This is super-interesting - may I ask where you learned about this part of the history of the postal service?
This article is way to fluffy to offer any proper advice.
If you're using auto-scaling or beanstalk or whatever, that's an example of AWS offering something where there isn't a norm, isn't a "normal way" to do it, and if you need that thing, your other alternative is to build it yourself - which is still an option should you decide to move away from AWS.
Lockin is avoidable switching cost which benefit the party imposing the switching cost. Are there examples of AWS doing that?
That said, I don't think they set it up this way to specifically cause lock-in, it's just technically difficult to allow users to manually manage replication alongside RDSs automated system.
Same with their cloud search - sooo obviously solr/lucene-based but 0 exposure to the lucene syntax.
Not a big deal though, it's not like it totally locks you in to use these.
Except it isn't. http://aws.amazon.com/cloudsearch/
"Amazon CloudSearch was created from the same A9 technology that powers search on Amazon.com."
Not using it (yet), but it's on our viewscope.
http://blog.mccrory.me/2010/12/07/data-gravity-in-the-clouds...