This blew my mind. What does this even mean in terms of logistics. How many people do you need to have just to add all those hard drives? How many new datacenters do you need to build every 5 years?
This blew my mind. What does this even mean in terms of logistics. How many people do you need to have just to add all those hard drives? How many new datacenters do you need to build every 5 years?
Of course, that's just YouTube, and Google has many other needs for data. But people forget just how big the denominators are on this quantity, and how effective Kryder's Law has been. They also forget how much reserve capacity there is in human labor; Google's datacenters have tiny employee counts because they are so automated, and could easily scale up into the exabyte/day range.
A more interesting question is what the differential rates of Kryder's Law vs. Moore's Law will do to how we architect software. Already, people in the know say that "disk is the new tape" - disk drive capacity has been increasing much faster than seek times, bus bandwidth, and available processing power, which means that you have to start treating the drive as a sequential storage device and not as a random-access platter. That's behind a lot of the shift from B-trees (as in conventional RDBMSes) to LSM-trees (as in BigTable/LevelDB), and also the resurgence of batch-processing frameworks like MapReduce. How does the software you build change when reading & writing data sequentially is really cheap, but accessing it randomly is expensive?
I thought we were already in that situation. Cache is king.
Of course at this scale you provision by prebuilt rack or even by container
It's unclear from the article whether YouTube's 1P/day is pre-replication or post-replication. I'd just assumed post-replication; it doesn't really change the conclusion in a material way. (Rather than the answer being "1 person", it becomes "2-3 people".)
And yes, you do need 2x (or more commonly, 3x) the disk space. You can use error-correcting codes to correct single-bit errors (Colossus uses Reed-Solomon), but that won't help you if there's a fiber cut and a whole datacenter goes offline.