Amazon’s Glacier secret: BDXL (2014)
storagemojo.com
storagemojo.com
The first comment on TFA says as much.
Edit: This is the actual Pelican paper: https://www.microsoft.com/en-us/research/wp-content/uploads/...
The systems described are servers that are over-provisioned with an order of magnitude more hard drives than you would typically see in a standard server. This allows a much higher ratio of total hard drive space to other components and cooling and power requirements, but at the cost of only being able to spin up some small fraction of the hard drives simultaneously. It then becomes a complicated scheduling problem of optimizing which hard drives to spin up and when based on the total workload of all the data you are trying to read and write at a given time. This system makes sense for data that is expected to be accessed on the order of once per year or less.
I wonder if microsoft going to publish specs (like opencompute?) to pelican servers?
It's probably just very widely striped, price-segmented data.
Also, see: https://news.ycombinator.com/item?id=4416065
When I was running exchange systems, our biggest challenge was delivering IOPS. We had to use SAN, and wasted significant storage because we'd spend our IOPS budget at 40-60% storage capacity.
I figured at their scale they would have similar problems.
Reading is pretty slow from glacier.
With exchange, we had all of this expensive, reliable SAN storage that would be perfect for a low requirement glacier like solution. Unfortunately, we lacked the ops mojo to pull it off.
for example I used to look after a quantum iScaler 24 drive robot, each drive was capable of kicking out ~100 megabytes a second. It was more than capable of saturating a 40 gig pipe.
However random IO was shite, it could take up to 20 minutes to get to random file. (Each tape is stored in a caddy of (from memory) 10 tapes, There is contention on the drives, and then spooling to the right place on the tape.)
Email is essentially random IO on a long tale. So, unless your users want a 20 minute delay in accessing last year's emails, I doubt its the right fit.
The same applies to Optical disk packs (although the spool time is much less.)
If Amazon has that problem with EBS, then selling that storage capacity as Glacier and using just the idle IOPS (or leaving a small bit reserved) allows them to sell capacity that would otherwise just be useless.
So if you use a lot of IO with Glacier, they are going to charge you like crazy, since you're potentially impacting EBS customers.
Tapes also match the slow retrieval speeds as you have to read the data out onto a drive linearly.
I do find it rather fascinating that AWS has managed to keep the technology used by Glacier, even at a high-level (i.e. disks vs. tape vs. optical), so under wraps. My personal guess is that it's powered-down disk drives on the grounds that's the simplest long-term solution but that's purely a guess.
Well, that basically tells you almost all you need to know. It's disk in JBODs. The only question is SMR vs conventional. Anyone who knows that can't tell you in public.
Also, I prefer privacy. To each his own, however.
How would you stop someone who gains physical access to your server?
( * ) So you don't need a power adapter.
I find that read only point in time backups gain value over time. Especially if you need to pull a file that would have been long rotated and replaced by newer backups on read write media (eg. HDDs).
Unfortunately the market for this use case is not large and this is greatly reflected in the prices and relatively hard to source high quality optical media. For BD this would be (inorganic) HTL Panasonic media which only has a market inside Japan itself. M-Disc is the other alternative although it only has proven itself within the DVD market, as classical HTL BD media is expected to be very similar in endurance to what M-Disc has on offer in the BD range.
Added bonus, 9 out of 10 customers actually preferred the feel of their data when it is restored.
That seems to have been the major stumbling block with higher capacity optical media, that one can't do the drag and drop writes that one have with spinning rust and flash chips.
Plus, if I was designing an archival system, it wouldn't be on blueray, unless there was a requirement for magnetic resistance.
In my experience, written CD and DVD only lasts for <10 years, if you're lucky. However studies show you can get 30-45, even 45+ years out of them.
Most Blu-Ray expectancy exceeds this due to the different non organic dye based layering and coating.
The one mentioned above, M-Disc was something developed for DARPA and supposed to last 1000 years in theory.
See:
http://loc.gov/preservation/resources/rt/NIST_LC_OpticalDisc...
http://www.zdnet.com/article/torture-testing-the-1000-year-d...
- Cheaper storage because data is heavily compressed
- Slow retrieval time due to slow decompression
However, is 2017, so we can say for sure that the extrapolation to 2016 in the linked article from 2010 was pretty good. It is too optimistic by a factor of ~2 in density, and ~10x in cost, but is spot-on even compared to most predictions from a year or two ago.
Since any whacko can claim they made a prediction in 2010, I double checked:
http://web.archive.org/web/20100322200343/http://www.storage...
Thanks for sharing the link!