Wayback Machine director outlines the scale of everyone's favorite archive
arstechnica.com
arstechnica.com
I donate $50 a year, because it's pretty much the only site I can check every month or so and reliably expect that it will be better since the last time I visited.
https://projects.propublica.org/nonprofits/organizations/943...
Looking at the form for ty2016, the most recent available, the Internet Archive's expenditures that year were about $16.4 million.
- donations (in-cash / in-kind) https://archive.org/about/credits.php
- digitization services https://archive.org/scanning
- web archiving services https://archive-it.org
src: I work @ Internet Archive on https://openlibrary.org
I suspect the bulk comes from a few generous individuals / corporations as this is not the most known platform.
Found this with a quick search: https://archive.org/about/credits.php
Then there is the running cost to power all those drives.
Also, when you have that many drives, you're going to be constantly replacing dead ones.
I have a feeling this is an expensive operation!
This is already factoring in 2x storage, so you just put each half in a different location.
> and they likely have some kind of offline archive too.
Okay so 2 million.
> Then there is the running cost to power all those drives.
Yes, that's one of many things that adds on top of the raw drive cost.
> Also, when you have that many drives, you're going to be constantly replacing dead ones.
It's not actually that bad! If you cram ten thousand drives into a few racks, your drive failures per week will be in the single digits. You could have someone come in once a month.
> I have a feeling this is an expensive operation!
It is, but my core point is that it would be expensive even if the drives were free. You need a certain amount of server horsepower and the physical resources to support it. This implies that you're not getting to stick 60 drives in a single chassis, so you need more labor as well.
There was an estimate last year of $1500 for IA to store and serve a terabyte of data indefinitely. I'd love to see a breakdown of that number, how much goes to day 1 hardware, how much goes to year 1 hosting, how long that hardware is expected to last. But even the day 1 cost is definitely a lot more than $50 for 2TB worth of hard drive.
My bad, I missed that!
> It is, but my core point is that it would be expensive even if the drives were free.
On reflection, and after reading your reply, I agree :)
> There was an estimate last year of $1500 for IA to store and serve a terabyte of data indefinitely.
I assume that's $1500 per year?
Total, not per year. Even S3 would only cost $250 a year to store a TB redundantly, and you could host 20 of them off a $200-per-month server or something.
Backblaze gets you a bit more than one copy for only $60 a year.
edit: Also this article mentions a bunch of things I didn't know about (e.g. playable 80s video game archive) so its worth reading.
Unsure if this is on RAID or ZFS.
Of course, you don't need to rebuild immediately. You're effectively working with (100,19), and you can put off the reconstruction as long as you like, maybe you don't reconstruct until someone wants to read the data or until enough other disks fail, and you can prioritize the I/O as low as you like. But in practice, super large encodings become more and more expensive as the size increases.
Since we have plenty of redundancy, let's keep things low-priority up to 3 drive failures, and try to rebuild each drive in 90 days. For bonus points, if two drives are rebuilding at once the increase in bandwidth is negligible.
8.5TB in 90 days is less than 10mbps. That means we could build servers with 50 drives and a single gigabit connection and if they were rebuilding every array at the same time it wouldn't even use half that bandwidth.
Real servers are going to have vastly faster connections and probably a lot fewer drives, so honestly that petabyte of I/O is not a big deal in context. In practice you could replace failed drives in a day, and the limiting factor is the speed of a single drive, not the network.
The arithmetic is correct but that's not the correct value. The I/O necessary for reconstructing one 8.5TB drive in a (100,20) Reed-Solomon group is 850 TB, which comes out to 875 Mbit/s, averaged over 90 days. It's not uncommon to see data centers with 10 Gbit/s connections, and sure, maybe it's a 10 Gbit/s per link on a Clos fabric but it's hard to claim that this bandwidth usage is trivial.
The point is not that the rebuild is prohibitively expensive or impossible, the point is that as the group size increases, the cost of data reconstruction increases and the cost of the encoding overhead decreases. At some point the cost savings from reduced encoding overhead are smaller than the additional I/O costs incurred by reconstruction. So the ideal encoding size is not as large as possible, but some medium size which balances the cost of the encoding overhead with the cost of reconstruction.
And consider that if any of this data is being served, you incur the 100x I/O penalty immediately.
> Real servers are going to have vastly faster connections and probably a lot fewer drives, so honestly that petabyte of I/O is not a big deal in context. In practice you could replace failed drives in a day, and the limiting factor is the speed of a single drive, not the network.
The Internet Archive has 24 disks per machine, I believe.
https://en.wikipedia.org/wiki/PetaBox
> In practice you could replace failed drives in a day, and the limiting factor is the speed of a single drive, not the network.
To rebuild an 8.5TB drive in 1 day requires 78 Gbit/s of bandwidth. Even inside a data center, oof.
You missed the part where I'm dividing the work over all the servers.
We're not having 119 disks read 8.5TB each and sending it all to the server with the new disk.
Instead, 119 servers are each recovering 85GB and then sending that to the server with the new disk.
Each server sends 8.5TB over the network, and receives 8.5TB over the network. You only need 10Mbps per server for slow mode.
> And consider that if any of this data is being served, you incur the 100x I/O penalty immediately.
1% of blocks take 100x the I/O. That's an average of only 2x. But because of data locality, you needed to read most of those blocks anyway. You might only have a few percent penalty.
> To rebuild an 8.5TB drive in 1 day requires 78 Gbit/s of bandwidth. Even inside a data center, oof.
It requires 78Gbps of switching bandwidth. It's easy to get a switch with terabits per second of total capacity. (less than $100 per port on eBay) https://i.dell.com/sites/csdocuments/Shared-Content_data-She...
So I assert that even 100ish drives per group is actually quite easy to handle.
> Instead, 119 servers are each recovering 85GB and then sending that to the server with the new disk.
Could you explain how this is possible? In other words, how do you arrange the blocks so that each individual device has enough data to reconstruct 85GB from another drive without doing network I/O? I'm not aware of a scheme where this is possible. Can you point to a paper or describe how this scheme works?
Simple arithmetic suggests that it is not possible to rebuild 100 different 8.5GB chunks from a single 8.5TB selected from an larger amount of data encoded with Reed-Solomon (100,20). The simple arithmetic suggests that you would only be able to reconstruct a total of 1.7TB from any individual 8.5TB disk, which is only 20% of what your scheme suggests. It is also unclear to me what would happen in this scheme if two disks were lost simultaneously, which is nearly guaranteed to happen at some point if you are playing around with 120 disks.
Simple erasure coding means that e.g. with (100,20) Reed-Solomon, you spread each stripe across 120 devices, and therefore need to read from 100 different devices to recover any one lost chunk from a single stripe. There are some schemes which reduce this but the gains are modest and there are tradeoffs involving e.g. encoding time or overhead, for example, the "Hitchhiker" schemes published by researchers at Facebook:
https://www.cs.cmu.edu/~nihars/publications/Hitchhiker_SIGCO...
Each server with a working drive reads a contiguous run of 117MB off its hard drive. It then sends the first MB to server 1, the second MB to server 2, the third MB to server 3... the last MB to server 117.
Each of those servers receives 117 chunks of a single 1MB stripe.
It then does a parity calculation to figure out the missing 3 1MB chunks, and sends them to the rebuilding servers.
We have now recovered 117MB of the disk. Our network data sent was 120MB per server. Every lost chunk was recovered via sibling chunks from 100+ different devices.
You can then optimize this by having the servers only send 100 chunks of each stripe across the network.
Once you repeat this operation 85470 times, sending 100+3MB each time, you have recovered 3 10TB disks with a network use of 8.8TB per server. (The efficiency is slightly worse than the 8.5TB for recovering one disk)
So you're still sending over a petabyte through your switch, but that's what good switches are built for, handling cross-traffic from every single port at the same time. A gigabit per server is acceptable, and 10gig approaches overkill.
And that really was all I was trying to say... you don't make encodings arbitrarily wide because the space savings must be weighed against the reconstruction costs.
I guess something to keep in mind is that you need bigger chunks to read efficiently, and smaller stripes to write efficiently. But in archival storage you don't care about write speed so the balance changes a lot.
You can make a sensible tradeoff for any given use case. For example you might get away with a 12,8 orthogonal nested encoding, that "only" has 67% space overhead and you will be able to rebuild a single lost member without reading the entire stripe.