14TB Hard Drives Now Available
anandtech.com
anandtech.com
They're not generally useful for the kinds of applications harddrives are classically used for (classic filesystems). They're mostly useful for applications that already use tape, with the benefit of much quicker random reads.
Log filesystems may be able to take advantage of them. There will necessarily be I/O overhead to garbage collect deleted files out of SMR regions. (Because you have to rewrite the entire SMR region to compact.)
Is anyone using or planning to use SMR drives in production? If you're able to share, I'd be curious to learn about your use case and how you plan to make efficient use of the disk.
There are ideas about "host-managed SMR" too where the HDD device presents the low level SMR device, and would need customized filesystems and drivers. But the idea isn't very attractive in the marketplace.
I run ZoL, and honestly they've surprisingly been relatively solid drives. You can expect roughly 10-15MB/sec or so write throughput per drive, and latency is of course pretty bad. Across 24 spindles though, I haven't had too many complaints - it's archival storage and reads are fast enough for most use cases - about 50-70MB/spindle sustained for large files.
I would not try to say rsync millions of tiny files against this pool - it would not hold up well. However for it's use case - write large files once/read occasionally I'm quite happy - I just wish price would have come down over the course of 2 years as I had originally expected. You can expect to pay $200-220/drive even today, and that's what I paid the first batch when starting my pool.
Out of 24 total spindles I had 2 early failures (within 120 days of install) but otherwise no other failures. These drives were replaced hassle free via RMA. My I/O pattern is pretty light - probably 100GB written per day across 24 spindles, and maybe 1TB read.
Basically if you try to do anything but streaming writes you're going to have a bad time. They are a bit more forgiving on the read side of the fence however. Don't expect these things to break any sort of speed records!
If I were buying today I wouldn't buy 8TB SMR - I'd pay the $20-30/spindle premium for standard drives. I'd have to look at the 14TB costs to see if the huge speed tradeoff would be worth it. When I first started using them, the cost per GB was compelling enough to give it a shot and I'm pretty happy with the results.
It seems ideal for a write-once/read-many media server.
Today I'm not sure the 8TB drives make any sense as prices have come down so far on regular 5400rpm slower drives that are much faster. These new 14TB spindles will be interesting to keep an eye on.
I imagine SMR won't really take off, if it does I'd expect more direct kernel/driver support for the hardware. Drive-managed SMR is always going to be exceedingly inefficient.
I could see a usecase for certain NAS / media storage purposes, but you'd have to not only use log structured filesystems, but also redesign the network protocols to support efficient writes, and possibly specialize block allocation to match the hardware constraints. You certainly wouldn't want say bittorrent writing directly to them, and even streaming-like services like MythTV may be a problem.
Of course there are some issues about balancing SSD size vs. cost and write cycles, but at least if designed properly the failure mode once the SSD part is weared down would be that it just functions as a regular (device managed) SMR drive.
I have three 8 TB Seagate SMR drives in RAID-Z (aka RAID-5) with a 500 GB Samsung 950 Pro M.2 as a L2ARC cache drive on the pool and it works beautifully for my workload.
In my experience, once the scratch area of the SMR drive is full, I get about 4-5 MB/s sustained write speed, which in my pool translates to 10 MB/s for the pool. Since the scratch area is about 25 GB, that means I can do 500 GB + 40-50GB of random writes before things slow to a crawl, and I have to wait (550 GB / 10 MB/s) = 55,000 s ~ 16 hours for the writes to flush.
I paid $179.99 each for these. Not Too Shabby.
Get 512 of them in a giant pool (especially if it's SMR aware) and the throughput issues start to be less of a concern for certain applications. Anything write intensive of course means you immediately look elsewhere.
Strangely enough, I could imagine it might work for a primary hard drive, as writes tend to be small and bursty allowing the SMR shuffling to catch up. But installation would take days.
Marketing them as "Archive" drives as Seagate did is the absolute wrong case for these. It's impossible to get any backup/archive copied over in any timely fashion. As a live mirror which gets piecemeal changes as they happen, then maybe. But that's still not an "archive".
I’d assume for these you should follow the same advice as SSDs, and issue the ATA Secure Erase command to have the drive wipe itself (or as is typically the case now, just it’s internal state and encryption keys):
You are spreading a lot of FUD based on one single anecdote with a used drive. Let me counter that anecdote with my own complete satisfaction of using such drive for over a year of daily backups without anything 'taking days'.
I agree! That's why I think "drive was used" should not be a disqualifier for seeing horrible performance. Why should a former filesystem matter?
> did you rewrite
Not my drive.
How do they calculate that? That works out as 285 years!
Half of these drives will last 285 years of constant operation. I find that hard to believe.
The methodology ought to be published, though.
Failure rate is also independent of lifetime. Drives have other measurements for lifetime.
Having 100 of this model would predict a failure about every three years (but that doesn’t mean it’d take 300 years to fail all of them). I’d be suspicious of a three year calculation, but it very well might turn out to be accurate. Remember that lifetime and failure rate are independent metrics. They’re warranted for five years, which is probably close to their predicted lifetime, and one drive of 100 failing in that time is certainly plausible.
Having 10 basically implies none of them will fail within a five-year lifetime. Again, plausible, but I think less likely.
Regarding longevity, often the predicted lifetime of a drive is close to its warranty. You will sometimes experience no issues exceeding design lifetime, and sometimes drives immediately explode. I’ve seen both, from four-year lifetime drives entering year 13 in continuous service to other drives buying the farm one day after lifetime and SMART wear indicator is fired.
As drives age, mechanical disruption becomes a much bigger deal. That rack of 13-year drives is probably one earthquake or heavy walker away from completely dead in every U. Even power loss, including from regular shutdown, will probably permanently end the drive when they’re far beyond lifetime. That’s the danger in a 24x7 server setting if you’re not monitoring SMART wear indicators (even if you are, really); power cycling your rack can, and does, trigger multiple hardware failures. All the time. If all the drives in it were from the same batch, installed at the same time, and an equal amount past lifetime, it’s very possible for the whole rack to fail when cycled — I have actually heard of this happening, once.
MTBF is unexpected failure. Design lifetime is expected failure.
There's always events dear boy, events!
MTBF lets you compare between two drives intended for a similar purpose.
For something that you own at home, the only medium that will last your lifetime in loosely controlled environmental conditions is paper.
Archival CDs might also last a lifetime (regular CDs certainly won't).
But I agree with your point that the MBTF of a single unit is not a way to predict it's lifetime.
It dawned on me when some older relatives died a few years ago that their papers, some very old, survive with a modicum of careful handling. My grandfather's immigration papers, great-grandmothers portrait on her wedding day, etc.
Today, many of us are one expired credit card away from losing all of that in digital form.
[1] https://en.wikipedia.org/wiki/Reliability_engineering [2] https://en.wikipedia.org/wiki/Survival_analysis
> Reinecke wondered if the host-managed SMR drives would actually sell. Petersen piled on, noting that the flash-device makers had made lots of requests for extra code to support their devices, but that eventually all of those requests disappeared when those types of devices didn't sell. Reinecke's conclusion was that it may not make a lot sense to try to make an existing filesystem work for host-managed SMR drives.
Edit: wmf points out that F2FS works reasonably well[0], until you have to do a garbage collection pass.
[0]: http://events.linuxfoundation.org/sites/events/files/slides/...
Looks like F2FS should work as of Linux 4.10.
I suppose what it does is similar to what drive-managed SMR drives do internally.
Helium is rare and critical for research and MRI machines it seems wasteful to use it in hard drives.
Note this is a much less wasteful use case than children’s balloons.
Also, modern MRI machines and other research equipment wouldn’t be feasible without massive amounts of storage to back them (though these drives are probably a poor fit for those use cases, since they are soft-realtime, and shingled drives can stall for a long time in the worst case).
They are terrible as single disks, or even in small arrays. Think of them as an alternative to tapes.