Lower Cost S3 Storage Option and Glacier Price Reduction
aws.amazon.com
aws.amazon.com
One thing concerns me however: Standard – IA has an availability SLA of 99%.
If this is just a reduced SLA but the actual availability is likely to be similar, that's fine. But if the actual availability is not expected to hit 99.9% -- say, if the backend implementation is "one copy in online storage, plus a backup in Glacier which gets retrieved if the online copy dies" that would be completely inadequate.
Hopefully we'll get more details over time.
> it’s vital that customers have immediate, instant access to any of [our photos] at a moment’s notice – even if they haven’t been viewed in years. [IA] offers the same high durability and performance ... so we can continue to deliver the same amazing experience for our customers.
The way this was phrased implies that this customer's use-case had a hard requirement that all of their data be in "online" storage at all times, and their satisfaction implies that IA does, in fact, hit this requirement.
I'm not sure what the 99% SLA means given that.
S3 Classic: Your file gets replicated on 3 live HDs (and/or HDs backed by RAID arrays—not sure about the internal S3 storage topology).
S3 Infrequent: Your file gets stored on 1 live HD (or single hardware redundancy component) and a copy in Glacier. If your live HD dies, your file will be automatically restored from Glacier to a new HD (but your data may be inaccessible during the automatic re-deploy).
Glacier: offline bluray combined with error correction accessed by robot arms, temporarily restored to live HDs on demand.
If a GET fails, just retry as usual (most higher-level libraries do this automatically, sometimes with a backoff mechanism).
Thanks! This is a very important detail which isn't documented anywhere: Retries are likely to succeed. A service where 1% of requests fail but failures are completely uncorrelated is far more usable than a service where 0.01% of requests fail but they keep on failing no matter how many times you retry them.
However, its SLA is a little more involved (for better or worse): http://aws.amazon.com/cloudfront/sla/
I'm not accessing S3 from EC2, though, so another benefit for me was it brought my S3 network costs way down.
I had to read these sentences a few times to understand what you were trying to say: "You now have the choice of three S3 storage classes (Standard, Standard – IA, and Glacier) that are designed to offer 99.999999999% (eleven nines) of durability. Standard – IA has an availability SLA of 99%."
No availability is mentioned for the others, but I assume it's 100%? Perhaps a simple table could help readers to scan and visually compare the values of two properties across three service classes?
Standard - IA: "Designed for 99.9% availability over a given year" [2]
Glacier: "Retrieval jobs typically complete within 3-5 hours." [3]
[1] https://aws.amazon.com/s3/storage-classes/#Amazon_S3_Standar...
[2] https://aws.amazon.com/s3/storage-classes/#Infrequent_Access
[3] http://aws.amazon.com/glacier/faqs/#How_can_I_retrieve_data
It does contradict the introductory blog post article here, but I'm assuming the actual documentation is more accurate.
When choosing between standard and this it would be helpful to understand the pros and cons. With the current description (below) it's as if the difference is only in pricing. But I assume there is a technical difference as well.
Also, the availability number could be explained better -- why is it different.
Standard - IA offers the high durability, throughput, and low latency
of Amazon S3 Standard, with a low per GB storage price and per GB retrieval fee.They might expect to be just as good as regular in the happy path, but are under-promising out of fear of some code-bug or other issue.
Either way, I'd definitely not migrate too soon for exactly the above reason.
Storage at 1c/GB/month, outgoing traffic at 1c/GB/month and no charge for incoming traffic. Data is replicated 3 times.
This isn't a joke: I can't find any documentation on the risk model that lets them estimate 11 9's and what class of risks it includes.
Glacier pricing is surprisingly complicated, and the actual
cost can be much higher than $0.01 per GB-month if you don't
read the fine print.
The biggest gotcha is that you can only access 0.17% of the data
you've stored in any given day without extra charges. So if you've
stored 1000 GB, you can only access 1.7 GB per day for free.
https://news.ycombinator.com/item?id=9184466 Glacier is only cost effective if you never
want to access that data ever again.
There are actually use cases like this, when you will almost certainly never want to restore this data, but just in case you put it in Glacier. For anything that you expect to ever reasonably want to restore in a reasonable timeframe (like a MySQL database backup), it just doesn't make sense: it's too slow and too expensive.And I'm not talking about regular backup data that is accessed quickly as needed. I'm talking about true DR that is only accessed when you have a catastrophic data center event.
Disaster recovery is being able to get your business up and running again. It would be pretty rare to have a disaster that wipes all your data, but hardware is intact and in perfect working condition. Your DR strategy needs to include hardware, location, people, everything.
The pricing is tricky (the per-GB price is cheap, but the retrieval can get horribly expensive). There's the fixed 4-hour delay for all actions (including listing stored files), which makes any interaction a pain. And there aren't really any good clients or high-level libraries that abstract away this complexity.
For a disaster recovery, I would certainly go for something simpler and easier to use. When everything is one fire, the last thing I need is dealing with a tricky API to restore the company files.
Look Glacier is great and the prices are really good. But it isn't something an SMB wants to be using directly. Now a large enterprise who can dedicate engineers to this, sure, but an SMB really wants to be utilising Glacier by means of a third party service in my opinion.
I think it is wise to think of Glacier as cold storage. So if you need recoveries RIGHT NOW, well, it may not be for you. If you can wait 24 hours? Sure (and, yes, I realise you can recover faster than that, but between transfer times, and actually starting the transfer it can take a while).
Glacier is PERFECT if you just need to restore a photo or document, and not the entire repo.
For better or for worse, Amazon fixed it, so we're still using Glacier.
We know that to be true for S3 Standard. Even S3 Reduced Redundancy claims "The RRS option stores objects on multiple devices across multiple facilities".
But I haven't seen Amazon make a similar diversity claim for Glacier. Perhaps it's implied by the "durability of 99.999999999% of objects" claim they make? Hard to achieve that durability in a single data center if there's a non-zero probability of something like a fire or other catastrophe.
S3 (and most AWS services) copies your data across multiple AZs but they are all pretty close together.
I used those words because the person I was responding to said:
$84 per year per TB is ridiculous
for geographically diverse storage.
The interesting question to me is "how close is too close"? Here in the Pacific Northwest we've had quite a few wildfires this summer. Even if Amazon's Oregon data center isn't anywhere near a forest, there are still failure modes that can affect a widespread area. For example, fires can disrupt power lines. They can also result in mandatory evacuations of large areas. They can also cause highways to be shutdown for many days. All of which can impact multiple AZs that are "close to each other".Glacier may be part of a 'scorched earth' disaster recovery, but it (almost definitively) can't be a primary option.
However, I find it interesting that in addition to the cost per GB to retrieve data, this new storage class also has a significantly higher per-request cost, too. Actually, it looks cheaper to upload an object as a different storage class and then transition it to Standard-IA, since PUT of IA costs $0.01 per 1000, but PUT of another class costs $0.005 per 1000, and the cost to transition another class to IA is $0.01 per 10000. It's a small difference ($0.04 per 10k objects), but if you store an obscene amount of data on S3, that seems like enough difference to matter.
Amazon, staring at its mighty armory, goes on the hunt for a tiny chink to repair.
Disclaimer: I work on Compute Engine (and not GCS or Nearline).
At $0.0125/GB/month, that means it costs $75 for 6TB per month.
But a 6TB hard drive costs less than $300, which means that assuming the data is stored on 3 hard drives for redundancy, they break even in less than a year.
However, hard drives seems to last at least 3-5 years on average, so this service seems to be priced at least 3-5 times as much as it costs to Amazon.
And there is even a $0.01/GB charge for retrieval on top.
There are other costs, but they should be relatively small at scale.
Am I missing something? If not, why doesn't anybody compete with Amazon and provide more reasonable pricing?
I That probably takes the margin from 80% of price to maybe 40%, in line with most retailers and whatnot
So you need to at least double the cost of the hard drive. Tripling the cost of hard drives might be a good assumption for S3/Glacier because you could assume they are building servers where hard drives make up 2/3 the cost of the server.
If you set up your own system, you have to manage all the technology and risk management strategies. Then you can see if you can do it cheaper than Amazon.
That's still speculation right? Some have theorized they use offline harddisks
So at least 6 HDDs, 1800$ of H/W and the monthly hosting costs for the 3 DCs and presumably 3 server chassis... Plus software to keep them in sync and you need to deal with the inevitable failures that will happen eventually.
Amazon guarantees 11 9's worth of durability (99.999999999%).
Backblaze sees between 2.41% and 7.77% annual failure rate on 6TB HDDs (source: https://www.backblaze.com/blog/best-hard-drive/) and they're the only ones publishing numbers like these over sensible number of HDDs.
* electricity
* labor
* bandwidth
* cooling
* servers, routers, wiring, and other infrastructure
* taxes on everything
* rent
* insurance
I have tried various glacier clients and they all seem to suck. So I have trouble tracking exactly what I have stored there sanely. :( Unusable.
For example, I have a cron job calling `s3cmd sync` for my photos on my iMac once a day.
It quickly becomes a nightmare for me. Hence I need git! http://natalian.org/2015/04/13/How_I_organise_my_media/
I think any active use of glacier is going to suck though. Just use that for archive.
and "infrequent access" is not on the price calculator
wtf is it so complicated to compare
https://aws.amazon.com/s3/pricing/
It has quite a lot of info.
Edit: you must view that page with JS enabled, otherwise no prices are shown. Perhaps that was your problem?
Firefox logs the dreaded
"Error: InvalidStateError: A mutation operation was attempted on a database that did not allow mutations."
which is known bad coding problem.