Google seeks new disks for data centers
googlecloudplatform.blogspot.com
googlecloudplatform.blogspot.com
This blew my mind. What does this even mean in terms of logistics. How many people do you need to have just to add all those hard drives? How many new datacenters do you need to build every 5 years?
Of course, that's just YouTube, and Google has many other needs for data. But people forget just how big the denominators are on this quantity, and how effective Kryder's Law has been. They also forget how much reserve capacity there is in human labor; Google's datacenters have tiny employee counts because they are so automated, and could easily scale up into the exabyte/day range.
A more interesting question is what the differential rates of Kryder's Law vs. Moore's Law will do to how we architect software. Already, people in the know say that "disk is the new tape" - disk drive capacity has been increasing much faster than seek times, bus bandwidth, and available processing power, which means that you have to start treating the drive as a sequential storage device and not as a random-access platter. That's behind a lot of the shift from B-trees (as in conventional RDBMSes) to LSM-trees (as in BigTable/LevelDB), and also the resurgence of batch-processing frameworks like MapReduce. How does the software you build change when reading & writing data sequentially is really cheap, but accessing it randomly is expensive?
I thought we were already in that situation. Cache is king.
Of course at this scale you provision by prebuilt rack or even by container
It's unclear from the article whether YouTube's 1P/day is pre-replication or post-replication. I'd just assumed post-replication; it doesn't really change the conclusion in a material way. (Rather than the answer being "1 person", it becomes "2-3 people".)
And yes, you do need 2x (or more commonly, 3x) the disk space. You can use error-correcting codes to correct single-bit errors (Colossus uses Reed-Solomon), but that won't help you if there's a fiber cut and a whole datacenter goes offline.
I've been thinking like this for a while and would tend to reduce the ERC (Error Recovery Control) to a minimum but the disks are still not designed to have these at very low numbers and Google has several interesting ideas in their paper along these lines.
It would be really awesome if they get the HDD makers to go along with this.
Somehow I don't think they will find things better than multi-million dollar research like HAMR for write-once read-many, but we'll see I guess.
I wonder if there will be 100TB spinning read/write drives by 2020, an exponential leap instead of incremental
Is there any reason for them to stay with 5.25 or 3.5 inch design for a datacenter?
Why not go back to 8 inch for massive surface area? Or is that too much mass to spin.
Though I wonder if the goals are slightly different, as Google is focusing on "Cloud Storage", rather than single hard performance.
Which as they mentioned, the Cloud doesn't need to have great reliability (since data is assumed to be already redundant), this may allow them to decrease the cost of hard drives for exchange in cheaper disks that may fail more.
Doesn't sound like they are expecting to increase size/performance alone with this effort.
http://www.sciencealert.com/this-new-5d-data-storage-disc-ca...
For continuous read/write workloads SSDs are the best for many reasons (energy efficiency, latency, IO bandwidth) and I think the world is moving that direction to replace spinning disks with these. I am hoping to see a bump in the SSD capacity as well. We are here now:
http://www.extremetech.com/computing/221303-the-worlds-bigge...
"An obvious question is why are we talking about spinning disks at all, rather than SSDs, which have higher IOPS and are the “future” of storage."
"The root reason is that the cost per GB remains too high, and more importantly that the growth rates in capacity/$ between disks and SSDs are relatively close (at least for SSDs that have sufficient numbers of program-erase cycles to use in data centers), so that cost will not change enough in the coming decade."
Cost per GB for SSDs is definitely catching up to HDD [1] and I think is expected to match or beat HDD within the decade. Power efficiency should be better on SSD and MTBF will be much better than HDD, so I would have thought TCO for SSD would have HDDs basically beat already.
[1] https://cms-images.idgesg.net/images/article/2015/12/ssd-vs-...
Which is a long time, but not totally out of bounds for the way we're currently using disks.
The coming decade is going to be very interesting, SSDs are pushing forward on an insane trajectory, which could either fall apart as the benefit of node shrinks disappear, or accelerate through a virtuous cycle as unit volume increases.
It's even plausible that spinning rust is on death's door in that time frame. The price floor of SSDs is much lower than that of hard drives (i.e. how much does an 8gb hard drive cost to manufacture today in volume? As much as a 500 gb drive), meaning that both mildly performance conscious and price sensitive consumers both move entirely over.
Enterprise customers will still chase the price/gb and storage density of traditional hard drives, but without the consumer dollars to invest in R&D, that could start to stall, quickly leaving them with diminishing advantage and obsolete in the blink of an eye.
The problem with spinning disks is that speed and reliability has not increased in line with capacity. There comes a point at which it doesn't make sense to make them any bigger (capacity-wise).
Just creating a filesystem on an 8TB disk takes hours. If you bring in disk (block level) encryption and the requirement to fill the disk with random data and then encrypt before creating the filesystem, you're looking at a multi-day task. Expanded to 100TB, you could be looking at a month just to bring a disk online.
On the reliability front - a spinning 8TB disk is probably about as reliable as a 1TB disk, so that means you have a 8x increase in probability of data loss, as well as ~ 8x more data to recover/re-distribute for every failure
That's what adding value looks like.
Are you sure you mean TB?
$ for i in $(seq 8); do truncate -s 1T 8tb.$i.img; losetup --find --show 8tb.$i.img; done
/dev/loop2
/dev/loop3
/dev/loop4
/dev/loop5
/dev/loop6
/dev/loop7
/dev/loop8
/dev/loop9
$ sudo mdadm --create --level linear -n 8 /dev/md8 /dev/loop{2..9}
mdadm: Defaulting to version 1.2 metadata
mdadm: array /dev/md8 started.
$ sudo time mke2fs /dev/md8
[lots of output snipped]
real 42m0.590s
user 0m2.470s
sys 1m0.440s
$ du -sh . # just for fun
65G .
That's certainly not fast, but it's not "hours" either. It's also a magnetic disk that was actively being used by other processes. Given how much data it wrote, it averaged 25MB/s, which is not really that fast. An 8TB SSD would be much faster.Am I being dense here, but why would you create a filesystem on the hard drive? The hard drive should only store file contents.
There's all kinds of seeking and syncing and random access needed for the filesystem. Seeks from file contents can't be avoided, but ones due to the filesystem can be.
If there's no metadata-on-ssd + data-on-disk filesystem for Linux, there should be.
A huge amount of work has gone into making ext4/zfs/xfs etc fast and reliable, plus you get other benefits like journaling, filesystem caching, fsync(), metadata caching (a la ZFS), so there are credible arguments for not going down the path of NIH.
I don't understand that argument, not if you think at scale.
You need to store 8000TB. You can run 8000 x 1TB, or 1000 x 8TB. You loose 10% of the disks per year (to keep the math simple). So after a year, you have either had to redistribute 800 x 1TB or 100 x 1TB, which is the same to a software defined distributed storage.
There are operational upsides to fewer disks, power usage, manpower, rack space etc. But of course more disks are faster.
Google wants to have more options than just bigger vs faster available. Since they know exactly, and can influence, how their software behaves, they can gain much from using specialized disks that balance things differently.
Xfs creates such a filesystem in seconds. As another reply showed ext4 is quite a bit slower since it preallocates the inodes; more modern designs allocate them on the fly as needed.
I'm imagining a 10 acre array of NAND
So the trick is to use every last IOPS that you can out of each and every disk, and that means provisioning and spreading out your data across an entire fleet of disks. The alternative, where you silo your storage in disk pools, and where some disks or disk pools have IOPS that are unused --- wasted, and where you make up for that by using flash for your hot and warm workloads is just a much more expensive way to do things.
It's for that reason that SMR drives aren't all that interesting. Sure, you get 20% more capacity, but at the cost of burning most of your IOPS for GC. Or if you move all of your SMR-friendly data to the SMR drives, that's cold data, and it means that you've wasted all of the IOPS that those SMR drives are capable of. Which means the rest of your disk fleet is now much hotter, and you might have to buy more CMR spindles to cover the cost --- at which point the cost advantage of SMR drives disappears.
(I helped to author the section about SMR drives, and why we are interested in hybrid SMR/CMR drives as a solution to this problem; read the paper for more details.)
One such request that is also mentioned in your article is to get more granular responses in the IO responses, in a SAS disk you can get correctable errors that do not hamper the behavior but will get you information on the internal status and actions of the disks. SATA has a problem with this since any IO Error in SATA will cancel all outstanding IOs. Though I think that there are device logs and such now that can be implemented to provide some semblance of a similar result.
> Flash is not getting cheaper faster than HDD's, once you control for write endurance.
Is the category of WORM-suitable data (like videos) not large enough to be interesting, or is there simply no WORM technology better than writable?I'm not aware of anyone selling standalone parts, but here is an IEEE paper: http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=124097... (512Mb in 2003!)
If Google really wanted to, they could make a couple of petabytes of slow-to-write memory by commissioning it themselves for low numbers of millions of dollars. But I suspect they value the rewritability.
Essentially the problem is that the HDDs and SSDs are built as black boxes and are not vertically integrated with a data center storage scheme. The major win will be when the HDDs and SSDs are slightly less self-sufficient and let the higher components in the stack manage them more in depth. Being able to tell the disk "do not bother with ecc for now" and try to get the data from your RAID and if that doesnt work come back to the disk and tell it "we really really want the data, spend as much as needed on this" is very strong in both getting the tail latency down and getting the reliability to stay at the same level.
For SSDs there is an added complication of garbage collection and the incredible amounts of over-provisioning needed to get consistent performance from them (at least 30% in an ME drive, the RI drives suck at consistent performance). If those can be reduced by more vertical integration that would reduce costs of the systems and make for better overall performance as well.
My most recent computer build used an M.2 slot SSD, and plugging that tiny little card into the motherboard seemed like the future. No cables required, and it was even smaller than the RAM sticks.
It does strike me though - why do disks do serial reads, couldn't the head read several tracks at once? Why not have several heads angularly spaced around the drive? Presumably such things have been tried and found lacking in some respect. Looks like I need to do some research ...
The paper in fact suggests this and also cites an attempt by Conner to introduce multiple actuator drives:
The thing that interests me more is the idea of more read/write heads. Add a second, independent set of heads for read/write and you will decrease latency a good amount. Or, a second dependent set of heads which would not provide as much of a performance boost, but would still improve latency a good amount.
> Why not go back to 8 inch for massive surface area? Or is that too much mass to spin.
One problem AFAIK is that larger disks wobble too much, more the faster they spin. Which is why if you break open a 3.5" 15k drive you'll see that the platters are more like 2.5", and a lot of wasted space. For 7.2k you can use bigger platters, but how much bigger? I doubt 8" is feasible.
Of course you can reduce the rpm further, but that has downsides too.
Assuming a compound annual growth rate of 25%, densities will be 10-fold larger in 10 years, and 100-fold larger in 20 years. (This actually happened in hard disks between 1972 and 1992.) But that would take us to 2035, not 2020.
At those densities, a bit will have to be stored inside a square about 7 or 8 nanometers on a side. I don't know whether that's possible.
Hmm, something up with the sums in the middle of that.
400 hours of video every minute is much more than one gigabyte per hour. It's way more than one terabyte per hour.
Working backwards:-
1 PB/day =~ 42 TB/hour =~ 728 GB/minute