15TB HDDs: Western Digital Unveils the Ultrastar DC HC620
anandtech.com
anandtech.com
That way they can be useful for more than archival.
Does that mean the drives simply will not work with an unaware filesystem or that it will work but performance will be poor?
However since I initially saw the news of these drives a few days ago Samsung also axed some Linux devs, which gives me pause and makes me reconsider the long term viability of this filesystem...
I believe Ceph now has abstracted the drives away through BlueStore, which simply puts a large RocksDB database on the drive, bypassing most of the functionality a filesystem offers. It should be much easier to make an SMR compatible version of the LSM-tree backend of RocksDB, than writing a full-blown file system.
RocksDB is one of several possible backends for object maps. There is a lot more to BlueStore than just omaps.
Also, BlueStore was actually designed with SMR drives in mind, however certain components of it are best placed on solid state media.
If you look at the design of, for example, DropBox Magic Pocket, or Infinidat & Qumulo, you'll notice that their HDD access is really as sequential as possible. And if your storage layer is thus already optimized towards sequential writes, why not take the opportunity and get some capacity "for free" by adopting SMR drives?
Seems like a decent use case. I'm not sure how well they'd work in a NAS with Raid-5/6 though. I'd been considering a new nas with 4-6 drives at 8-12TB already. Random I/O isn't my primary use, and I'm sure there are others at these sizes.
https://en.wikipedia.org/wiki/Nested_RAID_levels#RAID_50_(RA...
The correct answer is either a 3-mirror RAID1 or RAID6.
Bcachefs also promises some solution to this by allowing both erasure encoding and replication to co-exist, according to it's documentation.
However both raid 5 and 6 have 2 huge problems:
Data inflight at write time (power/hardware failures are more likely to corrupt the array, especially silently, which is the worst outcome).
Parity calculations require you to spin up the whole raid5/6 array during a rebuild, massively increasing the chance of a multi drive failure and a lost array. If one close-to-EOL drive dies, putting its sister drives through what is essentially an all day full tilt stress test is a terrible, terrible idea, and this idea keeps getting worse (takes longer) as drive sizes grow.
raid 0+1 sidesteps these issues mostly at a modest increase in drive count, its a no brainier for most setups.
How is that? RAID doesn't affect data persistence behavior in any meaningful way. FUA/SyncCache/etc are supported by RAID controllers same as the underlying disks in writeback enviroments, parity updates included. Put another way, if you FUA or flush the writeback cache, those operations won't complete in a properly implemented RAID environment until the data is persisted somewhere, even if that means passing FUA down to the underlying storage. Granted there are a number of ways to mess this up, RMW cycles in a controller that doesn't have some kind of persistent memory and flush on power restore. Anyway, none of this is any worse than what happens in any other WB cached storage technology.
Finally, all this fearmongering about loss on rebuild is also something that should be more fully explored in the context of the fact that decent RAID systems run background scrub operations on a regular basis. Those operations by themselves are going to "stress test" the array on a regular basis when its consistent and not degraded. I've actually got a fair amount of experience in this area, and I'm here to tell you that if you think this is a risk consider what happens to non-raided unscrubbed drives that have a lot of data silently bitrotting on the platters. That latter effect is nearly always the problem in RAID environments when someone starts a rebuild on drives/sectors that have been unread for extended periods of time. But, in the case of RAID, a properly implemented system won't fail a drive for a single read failure during a rebuild, instead reconstructing from the other drives and leaving the drive online long enough to complete the rebuild and then taking it offline.
Basically raid 1 setups don't actually fix any of these problems, except through the use of massive additional parity disks overhead. Overhead that can also be applied to other RAID algorithsm to much better effect. AKA a mirrored RAID 6 provides far more protection than a mirrored raid 0. Similar levels can be had with 6+6 in environments where that is possible, with trivial capacity overhead.
Battery and flash backup on controllers dosen't fix the problem of hardware failure (which is significant, especially on big hot controllers.
But much of this micro level redundancy is overkill as frequently one uses some kind of application level HA/redundancy as well. So, loss of a RAID5/6 disk in a single machine is the functional equivalent of loss of a any combination of RAID 0/1 in the same machine. You still need the higher level redundancy as well as a backup plan.
We could start breaking the discussion up into fabric attached vs direct attach RAID vs Software, but I think its sufficient to say, that RAID5/6 doesn't _increase_ the failure surface in any meaningful way when your not using fly-by-night RAID.
Edit: Maybe what your trying to say is that cache flush/FUA operations for a give piece of data don't cover the parity calculation and buffers? That is false, a controller should not be responding to FUA/etc until the entire (including the parity) block has been persisted. So if the controller dies during the operation the host OS is fully aware that the operation didn't complete. The given block is of course left in some unknown state in this case, but that is true of any write operation that fails like this, regardless of WT/WB/RAID/etc.
So even if you rebuild an array, a bad drive might've blown away all of your data already. If you were to compare this with ZFS' "raid" Z1 (same parity, different design) you get detection and protection against silent data corruption.
The rebuild isn't putting the disks under stress. The sister drive has already failed silently but you only notice this once you start the rebuild. The solution is to check the disks once a week by fully reading every sector.
"One" drives are the ones with a copy.
For archival purposes, though, you're probably better off with a normal RAID1 + some kind of JBOD setup (like with LVM); striping makes data recovery more difficult should you indeed lose all RAID1 sides of a given member.
Of course one should source RAID disks form 3 different vendors, to ensure that they are from different batches, and are not going to fail at approximately the same time.
In effect, you get the total bandwidth of (N-1) HDD's working in parallel. And the bandwidth of 100 HDD's doing sequential IO in parallel is really massive ( ~ 10 GB/s).
Examples of companies claiming to use this approach are Qumulo (rebuild in couple of hours), Infinidat (couple of 10's of minutes), ClusterStor GridRAID (now part of Seagate I think), or "Declustered RAID" in GPFS (IBM)
Thanks for pointing out that declustered/distributed rebuild RAID has many historical precedents (also 3PAR BTW) pre-CRUSH/Ceph.
I will note that you could easily build the MFT before you start transferring data. That's really your active 'dataset' here, and it's not very big.
I have a project where I routinely need to copy large amounts of data (3 to 8 TB) to a hard drive. Problem is, my files are all 512kb. So this is much slower than it could be...
If I write it as a single tar file I get excellent throughput, but the users who need to be able to work with the drive are unable to handle a tar file. They need to be able to plug the drive into a Windows computer and have it "just work".. which presents some problems.
Here are the ones for Q3 2018: https://www.backblaze.com/blog/2018-hard-drive-failure-rates...
Scroll down a bit and you'll see the annualized failure rates (AFR). The 10TB and 12TB ones seem to be pretty excellent.
https://www.tomshardware.com/news/wd-toshiba-hdd-hard-drive,...
And of course some people actually create the data themselves and don't download it.
I have 18 TB: 5 TB Seagate (x2), 4 TB WesternDigital, 2 TB WesternDigital, and 2 TB internal.
Backups take the most space - I fix laptops for friends from church, and they don't back up but still want their files to be safe. I had to shuffle some files around to free up 650 GB for a recent repair, mostly photos & videos.
Virtual machines use a lot of space too. I made VMWare Fusion images of every Mac OS version 10.5-10.13, Windows 95, 98, 2000, XP, 7, and 10, in several languages ( https://peterburk.github.com/i2018n ), and some Linux distros.
Another 1 TB is a dataset of Chinese characters from a machine learning project of mine ( https://blog.usejournal.com/making-of-a-chinese-characters-d... ).
Music, mostly from repaired iPods back in high school, accounts for a lot as well. There's some movies too, though I missed a chance to get 2 TB from a friend because I didn't have enough space at the time. If I upload those, even those that I legally ripped from CDs & DVDs, I'm worried that it'll trigger content filters.
For these, local disks are more useful than cloud services in my opinion.
You may find that there's a massive amount of data where you wouldn't expect it, such as in the Windows Temp directory - if so and it's a bunch of files named "cab_something", you can kill all of those and prevent recurrence with a little housekeeping.
Details: https://www.computerworld.com/article/3112358/microsoft-wind... (update log files in windows\logs\cbs get auto-compressed, but compression breaks and leaves big temp files if the file to compress >2GB)
[1] https://www.computerworld.com/article/3030642/data-storage/f...
I'm not an expert on transistor pitches, but here's a chart from Wikipedia for the 10nm - https://en.wikipedia.org/wiki/10_nanometer It's kind of impressive for HDDs considering that it's a 2 inch long mechanical arm that is able to move with that level of precision.
In every aspect, except price.
Samsung 1 TB SSD for 150€, Seagate 8 TB for 220€.
In fact at one point I made the mistake of enabling SSD caching on a NAS. The SSD became the bottleneck because of the limitation of SATA, ie one SSD on SATA is slower than 8 or 10 HD in RAID5. So unless you really need very high iops, HD are likely to be good enough.
Spinning drives are definitely not going away anytime soon unless there is a much more significant drop in the cost of SSDs.
Science investment requires a new technology to have a prospect of a return for most of the ~20 year patent lifespan for it to look like a good investment, and spinning bits of metal aren't that right now.
SSDs are higher performance than HDDs and have none of the packaging constraints. Flash storage is going to be put into everything and the economies of scale look quite good.
Storage is scaling but the r/w speeds of hdds aren't keeping up. Following the trend line and we see huge hdds that are functionally useless due to how long it takes to do disk operations.
HDDs only exist above tapes because of their performance. And only exist below SSDs due to cost. Tapes are the floor and SSDs are the quickly lowering ceiling. HDDs are likely to be crushed between.
That said, on the horizon of multiple years, I agree that the future scalability of NAND flash doesn't look quite as promising as HAMR/MAMR for hard drives. How that translates into actual product demand and adoption will probably depend on the relatively unexplored question of how much performance per TB our applications actually need. 40+ TB hard drives might not be fast enough to actually serve as nearline storage for that volume of data without eg. multi-actuator technology that essentially gives you more than one hard drive sharing a common spindle motor. Meanwhile, there's no question that QLC NAND flash definitely has adequate read latency and throughput.
With 100 read heads per platter, typical seek time is cut by a factor of 100. That won't let them overtake SSD's, but at least allow them to close the gap.
Going all the way to 100 read heads per platter would be insanely expensive and would massively increase drive failure rates, while still leaving them about four times slower for random reads than the slowest $35 SSD on the market. This will never turn into a viable product.
There have been many 'this will be the death of hard drives' technologies over the decades: zip drives, optical drives, tape drives (there was a time when they were predicted to be everywhere... never happened), CD (then DVD) writers, etc. Not to mention MRAM which has been the hottest tech that hasn't really happened yet for 3 decades. These were all going to be some combination of more durable and/or cheaper per Xb. But they all lacked the one critical advantage that hard drives had: massive economies of scale. Here's my prediction: spinning rust isn't going anywhere anytime soon.
I would agree on consumer systems (that often have one single boot drive), SSDs make a lot of sense. These days you tend to only see HDDs at the very low end.
From a personal perspective, most of my PCs are completely SSD. But I also have a media server (the largest storage space being reserved for MKV copies of my personal DVDs and Blu Rays). This is composed of RAID arrays of 6TB HDDs. Right now, cost wise, the highest "common" SSD is 4TB SSDs are roughly in the $800-$1200 range. 4TB HDDs by comparison are quite cheap, as low as $89 for a certain Seagate model (I used the Western Digital Reds which for 4TB at the moment are a little higher, $115... for the 6TB model it is $178).
While I definitely wouldn't say "HDDs will always be around", the price difference is very high right now to justify the superior properties of SSD. Backblaze seems to be in agreement here (https://www.backblaze.com/blog/ssd-vs-hdd-future-of-storage/). I guess the question is how long it will take for large capacity SSDs to scale down in price for them to be competitive. Until that happens, I imagine HDDs have a decent future left, if not on consumer devices then at least as drives for cloud / data center storage.
Lots of VPS and similar services offer SSD storage now, and I expect it to grow. The one I use didn't even a non-SSD option, and it was a pretty cheap provider.
SSDs offer orders of magnitude better random IO, I imagine that allows the hosting company to have more clients per storage unit, lowering the effective cost of SSDs.
Seagate and WD are spending mountains of cash to develop technologies like HAMR and MAMR which they expect to take them up to 40TB drives. These technologies require entirely new fabrication processes, etc. Very capital intensive.
Theoretically speaking there should be a point where the TCO of NAND Flash would cross HDD. But as NAND scales down, it also reduces its write cycle.
I was never a believer in HAMR, the technologies parent mentioned which Seagate announced in 2012, the idea just seems too unrealistic. The approach WD taking MAMR is much better using Magnetic Field. BPM is also far off, but BPM has been in research for nearly a decade, and MAMR should be here in 2019. All of these R&D will come to fruition in the next few years, where we expect HDD to scale to 100TB in the next 10 years.
https://blog.seagate.com/craftsman-ship/multi-actuator-techn...
That will close the performance gap a bit too I suspect. I have to wonder about power consumption though.
They accomplish this by essentially being multiple hard drives sharing a common spindle and helium-filled enclosure. As Seagate is currently implementing the idea, you still have only one head per platter, and at most one independently moving head per platter (but currently the stack of platters is just divided into two groups). Thus, sequential performance does not improve at all (and actually is reduced by the number of independent actuators), and random I/O increases by a small integer factor when the gap between hard drives and the slowest SSDs is already more than two orders of magnitude. However, power consumption shouldn't be much higher for this kind of multi-actuator hard drive over existing hard drive designs.
I looked at the drive actuators at the time, and I was incredulous when the guys told me that all operations were serial. I asked why they didn’t do parallel reads and writes, and I was told that technology was already common for mainframes but too expensive for consumer gear.
So, fast forward from 1989 to now, and I’m sure that idea will come back — sooner or later.
I'm sure eventually SSDs will be cheap enough that HDD will go the way of the floppy but we're not there yet.
6.25TB tape is $28 : https://www.amazon.com/HP-HEWC7976A-Ultrium-6-25TB-Cartridge...
6TB HDD is $120 : https://www.amazon.com/Seagate-Expansion-Desktop-External-ST...
EDIT: Commenter below points out this 6.25TB tape is actualy 2.5TB physical -- still cheaper, in terms of $/GB, but not the 4x I mention above -- closer to 2x.
LTO-6 (what you linked) is not 6.25TB, it’s 2.5TB, despite what Amazon says.
Then add the operational costs, which is the hard part, because the operational costs for a tape are very different from the operational costs for a hard disk.