Big leap for hard drive capacities: 32 TB HAMR drives due soon, 40tb on horizon
anandtech.com
anandtech.com
I didn’t have time to learn back when the first disaster hit with SMR drives got snuck in without warning and we had a small outrage event.
I’m happy for new drive tech but I’ve not seen much about the pros and cons of HAMR in real world use yet. Which could be because there’s no difference from traditional Perpendicular Magnetic Recording (PMR) hard drives, but then again it could have its own subtle set of trade offs entirely different to the ones that Shingled Magnetic Recording (SMR) drives do.
At their scale it’s easy to utilise these efficiently despite the limitations.
They’re not very good in a small NAS or disk array. You’d need a big write cache in front of them and an OS that natively understands the sector groupings.
I used to just rely on FreeNAS but it’s not as straightforward anymore. I’m having to consider Linux and looking at bCache FS and ZFS vs BTRFS and how all this compares on an all PCIe (m.2) flash drive setup… where a drive (or two) at this sort of size (~50TB) would make a great second backup copy of the flash array that can be started and stopped to periodically make the backups to save power… but then you have to think about copy efficiency since I don’t want to have them wasting read bandwidth from the flash drives… and the best is usually like to like file system copy ZFS -> ZFS and BTRFS -> BTRFS … so it adds another complication into a mix that is already far from simple.
So it’s become something I’m keeping an eye out for… hopefully before it becomes something I may purchase, someone will have already done a good writeup.
ZFS is a great improvement over traditional hardware RAID systems, but is still in many ways clearly a descendant of them. BTRFS has a slightly different mix of features from ZFS that make it a bit more flexible and a better choice for consumers who don't buy drives by the dozen. Neither has a great solution for caching/tiering with SSDs and hard drives. Ceph has a lot of features that would otherwise be almost exclusive to the hyperscalers, but is too complicated for something like a turnkey NAS. bcachefs aspires to eventually have most of the features you would want for a non-clustered storage system.
Zoned storage is something many of the above already have some degree of support for, but as a paradigm it has not even started showing up in the consumer computing ecosystem so many of the issues with adopting zoned storage for consumer systems aren't even being worked on.
Sounds like the only uncertainty is in what the next capacity past 24TB will be, but it'll definitely be SMR rather than PMR.
I’ve worked on these systems. You might as well be asking F1 drivers for tips to help you commute to work.
The hyperscale stuff is not built on top of ordinary filesystems. It’s all clusters of machines, and error correction is handled at the level of clusters. If you’re evaluating systems like ZFS and BTRFS, then you’re already working with a radically different tech stack.
At scale, your file metadata is stored in a distributed database of some kind, and the file contents are stored with forward error correction across multiple machines in a cluster. Or something similar, but stored across multiple clusters. The pricing models for cloud storage are designed around the usage patterns—like, if you know that some data is going to stick around for 90 days, then SMR is a win. If a file could get deleted at any time, then SMR is a loss.
Boston drivers: "hold my beer."
Most likely, cloud vendors have a large choice of sizes to try to match requested sizes of logical disks to physical disks, but if most users don't want / need disks that are this big (or bigger), then the provider will have a problem.
And there shouldn't be much of an impact on durability. The heating is very brief - a nanosecond or so. The idea is not to expand the material physically, it's to increase its magnetic permeability temporarily so that you can write the bit to it and then it becomes stable again.
https://en.wikipedia.org/wiki/Heat-assisted_magnetic_recordi...
Why does memory and other hardware tech always seem to magically come to fruition? At least compared to software
Is it because they're more conservative in their announcements? More certainty in a path forward? More $$$?
See also memristors, that HP in particular have been saying is coming in the next couple of years, fundamentally changing the entire computing industry, for decades.
FWIW, gallium arsenide was also supposed to take over from silicon for CPUs once speeds went above 25-MHz (or so). A senior colleague of mine said, "There are a lot of smart people whose kids's college tuition depend on making silicon a little better every year." And so, it came to pass. GaAs is definitely important but we've got multi-GHz silicon CPUs now.
[1] https://www.microsoft.com/en-us/research/project/hsd/ "Project HSD: Holographic Storage Device for the Cloud"
https://en.wikipedia.org/wiki/Bubble_memory
And then it just basically turned out to be a flop.
The neural net software algorithms have been around for decades. What made LLM’s feasible are the hardware advances to achieve unprecedented scale, just barely providing the ability (at great cost) to train today’s LLM models. Transformer architecture might be called a software innovation, but RWKV Raven gets similar performance to transformers and is built on decades-old RNN technology. So it is the hardware that was far more instrumental than the software in achieving LLM’s.
Counter to that argument: had google not done neural net research for google translate and proved the transformer approach scaled and performed well in their “attention is all you need” paper, people wouldn’t have spent the money to train foundation LLM models and we would not be having this discussion, so the software really mattered more than the hardware.
In reality I think it’s a little bit of both.
Even while they missed the opportunity, Gpt-1 was released in 2018 by OpenAI.
Then they incrementally added parameters in the next versions.
At what point does an influencial government say "it's unlikely there is substantial innovation to be uncovered for ___ product" and begin the process of deprivatisation for good of the public?
Most software (and hardware) is not like this.
Are there any good options right now for multi-tiered storage for the home lab?
LVM has writeback cache as an option, but I couldn't quite figure out how reliable that is, found some old posts with disturbing issues but not a lot of recent talk. Would also need to run ZFS on top of it for the ZFS features I rely on like snapshots and such, so feels like a Jenga tower solution.
I know Ceph has this option, but Ceph performance is abysmal for small installations from what I could benchmark.
Bcachefs looks like it'll be a winner, but it's still WIP so won't trust it with my precious data quite yet I think.
"tiering" in the traditional sense isn't even popular in the DC with things like Vast, Pure, etc all doing well.
Is there any reason you don't know the data set that is generally cold? The vast majority of home users basically have data that is trivially super cold that needs reasonable read performance - this is just HDDs. Then have a second mount point with things that you need greater performance, and have that be just SSDs.
https://openzfs.github.io/openzfs-docs/man/7/zpoolconcepts.7...
With large SSDs it seems alluring that one could have the majority of hot data on SSDs, essentially making the cold data even colder.
As you mention I could split my data in two, but that's a chore. If I need new data on the hot pool, I might have to first decide what got cold enough to move to the cold pool if there's not enough room left on the hot pool. That means I have to wait for that before I can save the new stuff on the hot pool. Searching for stuff means searching both places etc.
I forgot I stumbled over AutoTier[1] which could work, but it seems abandoned so again, not ideal for precious data. And it's FUSE based, so performance is likely not the best.
Your price of $10/TB might apply when HDDs are purchased in bulk, because at retail I see prices between $15/TB and $20/TB.
No HDD may be trusted to store data for more than 5 years and this duration is valid only for the more expensive models.
Assuming your price of $10/TB, one must buy at least a double HDD capacity than tape capacity, due to the short lifetime, so tape is at least 4 times cheaper. At the prices that I see at retail the difference is even greater.
The sequential transfer speed of tape is greater than that of HDDs, so archiving or retrieving many GB of data takes less time with tapes.
HDDs wear out and malfunction more frequently than tapes.
The only real disadvantage of tapes is the high cost of the tape drives, which makes tapes preferable only when more than 100 or 200 TB of data have to be stored.
After writing several hundred TB, I have achieved a decent money saving in comparison with using HDDs.
On the other hand, when someone needs to store only 100 TB or less, there is no chance to recover the cost of the tape drive, so tapes are inappropriate for such a case.
After writing several hundred TB
Sheesh, what are you writing!?The drive was super cheap, $150 used. I don't expect it to last forever, I plan to buy another tape drive as a backup and eventually upgrade to an LTO-6 drive.
After making 2 copies of all my important data, about 30TB on LTO-5 tape, I don't have that much to back up, maybe 2TB a month, but it's easy to justify buying a few more tapes every now and then. Buying 2 hard drives for redundancy is just not anywhere near as cheap as buying 2 LTO tapes for the same amount of data, even for people with less than 100TB of data.
The real disadvantage of tapes is what the drive may die at any time and if you don't have another drive then you can't recover. And a replacement drive wouldn't be cheap (new) or reliable (used).
When a HDDs dies, that is far worse than when a tape drive dies, because you not only lose some money, but you also lose data, which may be priceless, unless you have been careful to have backup copies.
While there are some companies that offer data recovery services from defective HDDs, for recent HDD models such services can be very expensive, comparable with the cost of a tape drive and much more expensive than the service of copying a good tape on a HDD, which can be done when it is not possible to buy a replacement drive immediately and the data is needed urgently.
It's same most of the time and LTO drive dies if you have at least any amount of dust while HDD doesn't give a fuck
> but you also lose data, which may be priceless
No difference if you sat on the tape with yours only one copy of your favourite porn.
And if you only have one copy of data then it doesn't really matter on which media it resides.
And oh, you CAN write the same data to two HDDs simultaneously with mirroring or just having a two copy jobs to a two separate hard drives, which would not only give you a physical separation but logical as well. For you to do the same with the tape - shell another $3000.
Yes they can fail mechanically (is this what you mean?) but you don’t necessarily lose your data.
Many of these bit flips, but not all, will be corrected when the sectors are read, due to the error-correcting codes that are used in HDDs.
This is not theory, I have stored data for several years on more than 60 HDDs of various capacities from both WD and Seagate, most of them being the more expensive models with extended warranty durations, but even so, only few of the HDDs did not have any non-correctable error after several years. (Fortunately I was careful to use redundancy, so there was no data loss.)
Moreover, some of the biggest HDDs that are available now are no longer suitable for long term data storage, because in order to improve the performance they store metadata in a flash memory, which has a more limited data retention time.
After more than 5 years the complete loss of a HDD should be expected at any time, but even after 2 or 3 years a few non-correctable errors are probable.
When a HDD fails mechanically, one might pay a data recovery service, but that might have a price similar to a new HDD, so if you plan to not replace your HDDs often enough with the hope of using data recovery, it is pretty much certain that the cost will be much higher than replacing any HDD preemptively when its warranty expires.
I do archive work and have 20+ discs from the 2010 era. Mostly the first generation of PMR drives. I have never had any data degradation problems.
You can also find lots of YouTube videos of people spinning up drives from the 80s and 90s which still hold their data without problem.
More scientifically, the phenomenon you talk about is modeled by the Arrhenius equation (1), where the activation energy to flip a grain is given by KuV/KbT, where Ku is the anisotropy of the magnetic media, V is the volume of a grain, Kb is the Boltzmann constant, and T is temp in Kelvin.
HDD manufacturers engineer this ratio to be >60 (usually targeting 70-90 to be safe). Media manufacturing is imperfect, so there is a log normal distribution of grains on real-world media, but if we assume that 60 is the energy barrier for all grains, a KuV/KbT of 60 would mean it takes 362 million years for half the grains to flip, assuming an attempt frequency of 10^10.
Where is my math wrong?
Assuming that your computed time is right, that means that there is a 50% probability that one bit of a HDD will flip after less than a week.
Most such bit errors will be corrected when a sector is read and the controller will rewrite a bad sector with a valid value, so the bit errors will not be cumulative in normal usage.
However when the data is stored for years without powering up the HDD, the bit flips will accumulate and they may pass the threshold needed to cause an non-correctable error.
While I do not remember to have ever seen non-correctable errors on the HDDs that I have been using daily, on identical HDDs that have been stored for years without being powered up I have frequently seen both cases when the drive reported non-correctable errors and cases when the drive reported no error but the file hashes used for error detection identified corrupted files.
The older HDDs with low data capacities had much longer lifetimes, but also the perception of those claiming that data has been stored OK on them may be wrong if they have not used any means to detect the corrupted files, because even if the HDD reports no errors, that is not good enough.
If you have a stored drive that is reporting errors, my starting assumption would be that something else is causing problems besides the platter—maybe the heads have gotten a bit of corrosion from humidity.
Still disagree?
Nevertheless, the experimental facts, both from my experience during many years with many HDDs and from the reports that I have read are:
1. Immediately after the warranty of a HDD expires, the probability of mechanical failure increases a lot. I have seen several cases of HDD failures a few months after the warranty expiration, while I have never seen a failure before that (on drives that had passed the initial acceptance tests after purchase; some drives have failed the initial tests and have been replaced by the vendor).
Therefore one should never plan to store data on HDDs beyond their warranty expiration.
2. When data is stored on HDDs that are powered down for several years, one should expect a few errors (I have seen e.g. about one error per 2 to 8 TB of data), which cause either non-correctable errors or wrong corrections that corrupt the data.
The effect of such errors can be easily mitigated by storing 2 copies of each data file on 2 different HDDs.
An alternative is to introduce a controlled data redundancy, e.g. of 5% or 10%, with a program like "par2create".
That works fine against wrongly corrected sectors, but when a non-correctable error is reported, many file copy programs fail to copy any good sector following a bad sector, so one may need to write a custom script that will seek through the corrupt file and copy the good sectors, in order to get enough data from which the original file can be reconstructed.
Storing everything on 2 HDDs, preferably of different models, is the safest method, as it also guards against the case when one HDD is completely lost due to a mechanical defect.
I've tested this on a small scale prototype level (one HDD and one NVME SSD) and there it worked quite well. But like I said it feels a bit fragile, lots of moving parts.
How are you expanding capacity? With traditional software raid I'd need to fail every drive and replace them one by one with identical higher capacity drives which requires several rebuilds from parity which means a long time and massive I/O loads and high risk of unrecoverable read errors just destroying the whole array in the process. It's easier to make a new array and move data over...
https://raid.wiki.kernel.org/index.php/Growing#Expanding_exi...
Btrfs has a more flexible block allocator but its parity implementation is still unreliable.
I used to do what you're doing. Last 6-7 years I've been running mirrored setup so I could expand by just adding a new vdev. Rebuilt it twice when I upgraded to significantly larger disks.
As I mentioned ZFS on top of LVM seems interesting in terms of features and flexibility. But reliability-wise... I'd like to run a more proper test setup to gain some experience before I'd trust it.
It will be more than 13 hours.
I guess if you want to be absolutely sure it's "clean" but most drives do intelligent remapping internally anyway, so "add and scrub afterwards" may be the way to go.
The cold data thesis hasn't worked out as well in the market as many storage vendors would have liked.
When the density of data on disks gets denser, it does so on both axes - x and y, or radial and tangential. But read/write heads only access one track at a time, so the speed-up is only along one axis.
Think about CDs, DVDs, and BDs. The discs can only be spun at a finite rate (somewhere around 10k RPM) before they shatter. The R/W head can read the data within any single track at full speed, but only one track at a time. From CDs to BDs, the amount of data per track increased (this doesn't affect the amount of time needed to read all the data from a disc), but also the number of tracks increased (this does affect the total time).
Also sequential workloads doesn't suffer from the write amplification and remain pretty high even if you are writing for hours.
So yes, a 32TB unit is very likely to debut at >$1k.
It would be nice if they could do a mini cartridge version ( 6 Folio Mini Disc inside a Cartridge ) like Zip drive with 2TB for consumer. Unless you are a Data Hoarder 99.9% of the consumer could have their own backup down with a few of these cartridge. And in theory they should last a lot longer than HDD or even SSD.
Really wish this isn't a pipe dream.
[1] https://arstechnica.com/gadgets/2007/08/new-dvd-sized-disc-t...
[2] https://www.storagereview.com/news/folio-photonics-working-o...
The etched plastic yes, but I think the reflective layer of CDs and DVDs is highly sensitive to humidity, shortening the lifespan to 10-20 years. You could probably recover it professionally, but not cheaply.
this is actually solved for HTL BD-R Blurays as they use inorganic material for storing data. Only the outer plastic is organic there. See https://en.wikipedia.org/wiki/Blu-ray_Disc_recordable
I'd wait until these actually ship, and we get some data on them. These end up basically being flaky new-age tape systems rather than HDD given SMR is required for this density.
Tape drives have multiple heads for IO, which allows some to fail along the way and you can still read/write your data. This sounds good - but it actually just means these tape drive heads are flaky. What do you do when too many fail? You call IBM or whoever and get them to replace your tape drive, or "reman" it by replacing the failed heads. This is the only way they actually achieve their warranties around lifetime read/writes.. they assume you'll fix the hardware along the way.
These HAMR drives have the same problem. The "heat assisted" just means they're using a laser to heat up a piece of gold, and sometimes this means the gold kinda drips around, and the head can be ruined. So their read/write lifetime numbers are pretty loose compared to PMR, and there is an assumption you'll "reman" these drives if they start to have failed heads for IO. However, the gold can even drip onto the platter, giving you permanent data loss anyways.
Lastly, they use SMR to get this density. SMR is not like PMR. PMR is what you think of with an HDD with many small blocks either 512B or 4KiB which you can read/write to. SMR has 256MiB (or 128) "zones" that you can only append, or reset. This means instead of being able to write randomly across the drive's capacity, you need to plan out your writes across appendable regions. This complicates your GC, compaction, and reduces your total system IO. Your random read performance is still better than Tape, but this basically turns your HDD solutions into something that looks a lot more like Tape. This is incredibly unpopular technology for this reason.
The market for these things is much smaller than most HDD vendors would like. You have a few hyperscalars that have figured out how to write out backups efficiently, but they want a lot of read IO, and reducing the spindle to byte ratio means you have less total system bandwidth to your data.
This means that the price per byte would actually need to be lower than their PMR drives, let alone wildly better warranty agreements, for these to be better TCO than current generation drives.
> The "heat assisted" just means they're using a laser to heat up a piece of gold, and sometimes this means the gold kinda drips around, and the head can be ruined.
It sounds like this tech is related to MO tech, which has been used for decades without issues (see: minidisc).
The laser is said to heat the spot on the platter to a bit over 400C, while the melting point of gold is over 1000C, so that doesn't jive. Did you meant something else?
> Lastly, they use SMR to get this density.
Why do they have to? SMR is just a way of fitting more data on a disc by overlapping tracks. I don't know why it is necessary to use this method on this technology.
Because they can. I saw a tech talk from a data recovery guy and he said some vendor (sadly doesn't remember which one) just have no new CMR drives at all.
I have a 5 TB external drive I bought many years ago, and despite its big physical footprint, the new drives still don't feel like a justifiable upgrade, given their elevated price for a small increment in storage.
Your move, NSA.
The practical issues of RAID rebuilds also don't quite work that way - the "drives die while rebuilding problem" has two components:
(1) is that beyond a certain size, the risk of bit-error while resilvering becomes a certainty due to all the data you're reading - ZFS fixes this problem with it's checksums basically.
(2) If all your disks were installed at the same time, then the chance of another disk failure is high because they're likely all in the same part of the mean-time-to-failure distribution.
(3) the risk of an error is higher because you're doing a lot more activity then normal, and so the additional stress is increasing the chance of triggering a failure.
Though if you want specific numbers, what would you estimate as the per-drive odds of 8TB and 40TB drives dying during a resilver? Let's say they take 2 and 7 days respectively.
But I would also expect these drives to be cheaper per TB, at least by the time HAMR is 2-3 generations old.
If you only do 5+3 to keep the same 160T available, it's 256TB or 128% paid for and the same 100% usable, so an even steeper 22% discount to break even.
And nobody that's price conscious would make a 5+3 RAID so I'm not very worried about such an extreme scenario.
On a related note, remirroring or rebuilding a 32TB RAID drive would probably take days.
Local drives are more fast cache for data that should be stored reliably elsewhere (maybe hot, maybe cold, depending on provider and cost concerns).
1. https://www.servethehome.com/discussing-low-wd-red-pro-nas-h...
S is Shingled, H is Heat Assisted, while the article mentions P for Perpendicular which became conventional after nobody has room for L or Longitudinal.
Explainer: https://www.linuxadictos.com/en/diferencias-entre-smr-cmr-pm...
WD NAS apology tour: https://blog.westerndigital.com/wd-red-nas-drives/
Seagate's CMR/SMR list: https://www.seagate.com/products/cmr-smr-list/
Maybe file fixity is in the future, where the concept is more foreground than before.
https://blog.seagate.com/craftsman-ship/hamr-next-leap-forwa... > Power, heat, and the reliability of related systems is equally nominal. HAMR heads integrated in customer systems consume under 200mW power while writing — a tiny percentage of the total 8W power a drive uses during random write, and easily maintaining a total power consumption equivalent to standard drives.