The pricing is tricky (the per-GB price is cheap, but the retrieval can get horribly expensive). There's the fixed 4-hour delay for all actions (including listing stored files), which makes any interaction a pain. And there aren't really any good clients or high-level libraries that abstract away this complexity.
For a disaster recovery, I would certainly go for something simpler and easier to use. When everything is one fire, the last thing I need is dealing with a tricky API to restore the company files.
Look Glacier is great and the prices are really good. But it isn't something an SMB wants to be using directly. Now a large enterprise who can dedicate engineers to this, sure, but an SMB really wants to be utilising Glacier by means of a third party service in my opinion.
I think it is wise to think of Glacier as cold storage. So if you need recoveries RIGHT NOW, well, it may not be for you. If you can wait 24 hours? Sure (and, yes, I realise you can recover faster than that, but between transfer times, and actually starting the transfer it can take a while).
Glacier is PERFECT if you just need to restore a photo or document, and not the entire repo.
For better or for worse, Amazon fixed it, so we're still using Glacier.
Glacier pricing is surprisingly complicated, and the actual
cost can be much higher than $0.01 per GB-month if you don't
read the fine print.
The biggest gotcha is that you can only access 0.17% of the data
you've stored in any given day without extra charges. So if you've
stored 1000 GB, you can only access 1.7 GB per day for free.
https://news.ycombinator.com/item?id=9184466 Glacier is only cost effective if you never
want to access that data ever again.
There are actually use cases like this, when you will almost certainly never want to restore this data, but just in case you put it in Glacier. For anything that you expect to ever reasonably want to restore in a reasonable timeframe (like a MySQL database backup), it just doesn't make sense: it's too slow and too expensive.Disaster recovery is being able to get your business up and running again. It would be pretty rare to have a disaster that wipes all your data, but hardware is intact and in perfect working condition. Your DR strategy needs to include hardware, location, people, everything.
Glacier may be part of a 'scorched earth' disaster recovery, but it (almost definitively) can't be a primary option.
And I'm not talking about regular backup data that is accessed quickly as needed. I'm talking about true DR that is only accessed when you have a catastrophic data center event.
We know that to be true for S3 Standard. Even S3 Reduced Redundancy claims "The RRS option stores objects on multiple devices across multiple facilities".
But I haven't seen Amazon make a similar diversity claim for Glacier. Perhaps it's implied by the "durability of 99.999999999% of objects" claim they make? Hard to achieve that durability in a single data center if there's a non-zero probability of something like a fire or other catastrophe.
S3 (and most AWS services) copies your data across multiple AZs but they are all pretty close together.
I used those words because the person I was responding to said:
$84 per year per TB is ridiculous
for geographically diverse storage.
The interesting question to me is "how close is too close"? Here in the Pacific Northwest we've had quite a few wildfires this summer. Even if Amazon's Oregon data center isn't anywhere near a forest, there are still failure modes that can affect a widespread area. For example, fires can disrupt power lines. They can also result in mandatory evacuations of large areas. They can also cause highways to be shutdown for many days. All of which can impact multiple AZs that are "close to each other".