SMR does not have a great reputation, to put it in drives specifically targeted for NASes and enthusiasts is a very bold move.
Classical economics likes to think of firms as rational entities working to maximize profit over the long term. In my experience that makes much of what companies do inexplicable. Instead I view firms as large groups of individuals all acting in their own self interest, this seems more complex but really all you have to know is what the incentive structure is to understand the choices a company makes. Similarly I've found you can also do this in reverse. It is a virtual certainty that from a compensation perspective moving to SMR was better for the individuals who made that choice.
Managerial and behaviour theories here:
https://en.wikipedia.org/wiki/Theory_of_the_firm
Linked to from the Economics article under the Theory section:
https://en.wikipedia.org/wiki/Economics
The theory your suggesting here has been known since at least the late 50s.
You might also be aware of, or find interesting anyway, the principal–agent problem:
https://en.wikipedia.org/wiki/Principal%E2%80%93agent_proble...
Edit: typo
Transaction costs, agent problems, information economics, institutional economics, mechanism design and so on are powerful insights into the world around us. I've learned more about how software grows and spreads from economists than from software engineers.
I can see how it might have come across as brusque, I certainly didn’t intend any insult.
SMR is fine if you know what SMR is and what you're doing. It might even be fine given a sufficiently advanced translation layer (which apparently doesn't exist yet). The problem isn't SMR, it's selling a product which uses SMR and hiding that fact from the consumer, and I can't forgive WD for that.
https://www.snia.org/sites/default/files/SDC15_presentations...
You put the journal and metadata on the continuous part and data on the shingled part, making sure you only ever add to what was written previously in a shingled segment by using copy-on-write. And sporadically you reclaim a segment.
This should be much more efficient than a block layer translation layer that knows nothing about the filesystem.
Will my hard drive's workload ever resemble that of a NAS RAID rebuild? Or will there be another workload which breaks SMR drives?
Key here is that workloads change, and weird issues like this one lead to compounding failures in emergencies. A system might have an SMR-friendly workload during it's normal operations, but during a cyberattack? When debugging an outage? It's impossible to predict what I'll be doing when the unexpected comes up.
Above all else, I want my storage medium to be reliable and to not have issues like this. It's clear SMR drives aren't there yet.
Yes, if there was a sufficiently advanced translation layer, or the right abstractions to the OS, SMR might be fine. But the present state-of-the-art of SMR technology is very obviously unsuitable for any application where you care about your data.
I'll mention early SSDs were in a similar state many years ago, with wear leveling algorithms. But there was a difference in the level of transparency there.
Apparently f2fs does well with it?
The real problem comes when you can't work like that, in particular this happens when you're low on free space or if you're in a situation where relocating data isn't possible or known that it needs to happen. RAID setups are typically the worst case for this, since they'll effectively treat the entire disk as used and can't relocate data on it. ZFS is a bit more complicated with this but effectively ends up the same.
Eventually someone listens to that person, and decades of painstakingly built-up brand value gets thrown out the window in order to bring home six nickels tomorrow instead of five.
I get that the reds are more performance oriented, but capacity is still a key factor (if not, you just go SSD and be done with it).
Magnetic disks are valued on capacity only because people considered CMR HDDs a performance floor. When they realize how bad SMR is, I think performance becomes a factor.
If you're talking about GB per volume, SSDs beat HDDs of all stripes there: the biggest HDDs are 16 TB in a 3.5" form factor, while the biggest SSDs are 100 TB in the same form factor. 10 or 50% improvement in HDDs wouldn't begin to bridge that gap.
SMR should be understood to be sort of a fast Tape, rather than slow CMR HDD.
In fairness, more like an array of tape drives. ;-)
I get what you're saying about how it hasn't panned out in reality, but the question was why did they even try SMR, and I expect that was motivated by the misguided potential of the win.
I wouldn't go that far. If I can assume that the firmware won't cause a lot of problems all by itself, I think in my desktop I'd slightly prefer a 7200RPM SMR drive to an equal-price 5400RPM CMR drive.
And no matter what kind of hard drive(s) you're using, take a $20 SSD and use half of it as a cache with writeback enabled.
I suspect Blue and Red drive have a lot of common in manufacturing they switched all drives in same capacity together.
It was well known before that Red are trash compared to Red Pro which IIRC are rebranded HGST DeskStar NAS.
If you had 32GB of Optane and 4 TB of SMR properly tuned that would be one heck of a desktop drive, but (1) you have to really eliminate all the bottlenecks and (2) who needs a desktop drive large enough that SMR is worth it?
When that happens depending on what your doing, the drive might just get kicked from an array, or your whole system might freeze if some critical page needs to be read back in and it can't get sufficient priority and a command slot.
I guess the vast number of people won't ever write more than a buffers full of data, or the drive will never get fragmented sufficiently that even small write operations amplify into entire drive rewrites.....
For that to work, you need the host OS to be able to see that the drive has just "shrunk". Obviously, you still have data on it while it's shrinking, so the reality is the OS needs to give the drive a list of sectors which can be "shrunk away".
One simple approach would be for a daemon to create a massive file filled with "MAGICBYTESMAGICBYTESMAGICBYTES...". As that data is 'written' to the drive, the drive sees that it's the magic bytes, and rather than storing the data, simply marks those sectors as no longer needed. As soon as enough sectors in a row have that designation, reformat them as non-smr, and use as regular (non-smr) disk space.
Then, a few hours or days later, the drive can rewrite the data back to be SMR, then tell the daemon, which can remove the magic bytes and delete the files, and your free disk space increases again.
The failure mode of this is that your "6TB" SMR disk ends up with a 3TB file filled with MAGICBYTES, and 3TB of your data. But for the same money, you'd only get to store 3TB of data on a non-SMR drive anyway...
You could nearly get there today with drive firmware only, and no special OS support, using TRIM. The OS can already tell the firmware what space it doesn't need, so the firmware could do what you suggest with this information. The only catch (and difference from your proposal) is that there's no way for the firmware to tell the OS that the space is currently not available, so if the OS puts pressure on that space the firmware would still need to block writes while it SMR-izes and frees the space it previously borrowed.
In other words, drive firmware could use space freed by TRIM for SMR-ization caching today, with no OS modifications.
Shrinking the disk is _MUCH_ harder, there are various enterprise storage arrays which are basically thin provisioned dedupe/etc arrays, and they overwhelmingly just lie about the capacity and throw up big warnings if the physical capacity is being approached. Then depending on which filesystem your running in linux, if they abort writes, there is a good chance the filesystem is damaged (some handle it better than others, and its getting better).
Here is their conclusion:
-------- begin quote
We want to be very clear: we agree with Seagate's Greg Belloni, who stated on the company's behalf that they "do not recommend SMR for NAS applications." At absolute best, SMR disks underperform significantly in comparison to CMR disks; at their worst, they can fall flat on their face so badly that they may be mistakenly detected as failed hardware.
With that said, we can see why Western Digital believed, after what we assume was a considerable amount of laboratory testing, that their disks would be "OK" for typical NAS usage. Although obviously slower than their Ironwolf competitors, they performed adequately both for conventional RAID rebuilds and for typical day-to-day NAS file-sharing workloads.
We were genuinely impressed with how well the firmware adapted itself to most workloads—this is a clear example of RFC 1925 2.(3) in action, but the thrust does appear sufficient to the purpose. Unfortunately, it would appear that Western Digital did not test ZFS, which a substantial minority of their customer base depends upon.
These tests may not be great news for either the American or Canadian class-action lawsuits currently underway against Western Digital, but they aren't the end of the line for those lawsuits, either. Even in the best case, the SMR models of WD Red underperform their earlier, non-SMR counterparts substantially—and consumers were not given clear notice of the downgrade.
If the same firmware was being used to make substantially larger drives available to consumers than would otherwise be possible, and the limitations of those drives were adequately explained, we would probably be gushing over its utility and function. Unfortunately, Western Digital has so far only chosen to use it to cut manufacturing costs on small disks, without even passing the savings along to the consumer.
-------- end quote
That last paragraph may be key. It sounds like if these drives had been marketed a little different and priced a little different they could have been a good deal for many NAS users.
I wonder what the chances are that they were originally intended for just that, and something got mixed up or miscommunication between design and release?
[1] https://arstechnica.com/gadgets/2020/06/western-digitals-smr...
After we published our numbers, we found that it was not just ZFS, but other NAS vendors as well. People were sending us their experiences on non-ZFS systems. For example, Synology users were having issues and Synology took them off their compatible list. QNAP for its part is working on ZFS as we discuss in there as well.
The new 4 drives were the infamous 6TB EFAX. The phase 1 took 12 days, the phase 2 took 3 days. So my personal experience is radically different than Ars Technica's.
https://www.backblaze.com/blog/hard-drive-stats-q2-2019/
https://www.backblaze.com/blog/backblaze-hard-drive-stats-q1...
EDIT: For those who can't use the Internet to look things up, Sandisk and G-Technology and some random thing called Upthere are all WD brands.
https://www.theverge.com/2016/5/12/11662018/western-digital-...
It's interesting that Toshiba apparently acquired Hitachi's production assets during the HGST→WD acquisition. HGST had a pretty good reputation at the time, so I feel like this bodes well for Hitachi if their drives are indeed a continuation of that.
If I do need some spinning rust at some point I guess I'll have to give 'em a whirl.
1000000 = 1M
Can I store just ONE byte at a time? Can I expand by singular bytes in a normal system?
For hard drives, and even flash memory, the physical size of a sector converged, rather naturally, to a multiple of system word size, and since that had also converged to a power of 2, to a larger power of 2.
Thus hard drives have sectors that were 512 bytes (4096 or '4Kbits'), and also why the smallest write available for modern drives is a multiple of 2.
Binary is the native nomenclature of computers. It only makes sense that related equipment should also use the native (binary) engineering unit approximations as a result.
Though for communications equipment, they do have variable sized packets, which even though computers generate 8 bit 'octet' based signals might technically be used by some other format that doesn't; and talking about the actual speed on the wire conveying that in maximum possible raw bit states is a tiny bit confusing when the rest are binary units, but it's less strongly wrong than a system where the lowest addressable (real) unit is binary based.
We know how well that worked out historically, right? e.g. https://en.wikipedia.org/wiki/Ounce "the international troy ounce is equal to exactly 31.1034768 grams", "The international avoirdupois ounce is defined as exactly 28.349523125 g", "A fluid ounce (abbreviated fl oz, fl. oz. or oz. fl.) is a unit of volume equal to about 28.4 ml in the imperial system or about 29.6 ml in the US system".
By not respecting the definition that kilo = 1000, you are deliberately sowing confusion and undermining the metric system as a simple, consistent, universal measurement system.
The good news is that I've basically stopped buying spinning rust, so this is more of a "well bummer" type of mood than an "oh shit I gotta scramble for a better primary vendor" mood. Even for server/NAS use I'm more likely to go for a Crucial/Kingston mix nowadays than I am for a WD/Seagate mix, both because SSDs have matured enough to be more reliable for a lot of things (less moving parts = less mechanical wear/tear and less sensitivity to vibrations¹, and with RAID a dying flash module is less of a "game over" moment - though things have gotten a lot better on this front, too) and because they perform a lot better (both in terms of speed and in terms of thermal/energy efficiency).
----
For example, consider two different 2TB cell setups:
- SSD: Crucial MX500 + Samsung 860 QVO = $473.73 on Newegg
- HDD: WD Blue WD20SPZX + Toshiba L200 = $168.57 on Newegg
On the one hand, the SSD cell is a whole 2.81× more expensive. On the other hand, the SSD cell is only 2.81× more expensive.
This narrows to 2× in the case of a 1TB (well, technically 960GB in the SSD case) cell setup:
- SSD: Crucial MX500 + Kingston A400 = $236.54 on Newegg (this is my typical setup for a RAID1)
- HDD: WD Blue WD10SPZX + Toshiba L200 = $117.13 on Newegg
Now, to be fair, this is excluding Seagate, which is stupidly cheap and would widen that gap (and historically I've bought Seagate strictly as a secondary for these sorts of "cells", but given my bad experiences with Seagate I'm inclined to ignore them and go with WD+Toshiba if possible for a given target size). It's also excluding 3.5" drives (and therefore notably excluding WD Reds entirely), which doesn't help for a 1TB cell (Seagate Cheetah NS + WD Red = $118.75) or a 2TB cell (Seagate IronWolf + WD Red = $167.74), but does help for bigger cell sizes that wouldn't be possible with 2.5" at all without major sacrifices in vendor redundancy and price (for example, a 4TB cell w/ Toshiba N300 + WD Red = $225.91, v. Samsung 860 QVO v. WD Blue = $1049.17 - even higher if I wanted to keep Kingston in the mix - or two of the above 2TB cells for only slightly cheaper).
On the other hand, this is including 5400 RPM drives; if I required 7200 RPM (as most NAS drives are) then the 1TB cell cost jumps up to $134.98 (WD Black + Seagate BarraCuda), and if I needed 10k RPM or SAS then I'd probably be springing for a WD XE WD9001BKHG + Toshiba AL13SEB900, which would bring the cell cost up even higher to $164.75 (while also losing 100GB of capacity), or even higher to $186.38 if I wanted to swap the WD for a Seagate Savvio ST9900705SS and have that 64MB of cache on both sides of the mirror (thus bringing the cost difference to a mere 1.27×). I'm also artificially keeping out a whole swath of SSD manufacturers in the above figures; never heard of "Team Group" or "Goldenfir", but if I had zero qualms about reliability I could put together a 1TB cell with those for $182.98 (on par w/ the 10k RPM cell), or for a 2TB cell I could go with a Crucial BX500 + Patriot P200 for $399.98.
----
That all being to say: SSDs are a lot cheaper nowadays than they were even a couple years ago, and at typical consumer scales they're plenty viable; "good enough" capacity, while also having better performance and lower power consumption.
Let's bump this up a notch. WD RED 8 TB drives are 213 now on newegg (base price). Two of those quadruple the SSD storage for the same price. You're now reaching 4x prices. If you go the data hoarder approach and shuck 8TB drives, that drops it to 290 for to 8 TB drives, putting you at 60% of the cost for 4 times the storage.
And that does require 3.5" drives (once you go to 2.5" that 4× gap shrinks to 2.7× for 5400RPM or 2.1× for 7200RPM+ when going by pure price/GB; per PCPartPicker, a 2TB 5400RPM Toshiba MQ04ABD200 is at 3.5¢/GB and a 7200RPM 1TB Seagate Constellation.2 is at 4.6¢/GB, while a Crucial BX500 is at 9.5¢/GB, all three of these being the cheapest per GB in their respective groups).
And I know the inevitable response to that is 'well real data hoarders are fine with 3.5" drives and/or 5400RPM', but "real" data hoarders also might appreciate being able to saturate their gigabit Internet connections with the data they're hoarding, or being able to hoard that data more compactly and quietly (I can cram 16 2.5-inchers in a 2U SuperChassis and still have room for a 5.25" drive, or 24 without) and without having to worry nearly as much about vibration or heat killing the drives (and thus the data on 'em).
----
Speaking of a 5.25" drive, if price/GB is really all-important, it seems like tape would be the ideal option, no? A bit high of an upfront cost (an HP EH957A LTO Ultrium 5 tape drive runs for $801.15 on Newegg, though there are some much cheaper used LTO 5s on eBay), but at $18 per 1.5TB cartridge (going with the IBM one) it's hard to beat that 1.2¢/GB. As soon as you cross the 176TB mark (if not earlier; my napkin math is equating 9TB of tapes to a shucked drive), tapes would be cheaper per GB even factoring in that brand-new drive (for retail Reds, the crossover point would be 64TB).
They say tape's dead, but it looks like it's still got an edge there, even with older tech. A tiered SSD + LTO 5 approach seems like it'd be the best of both worlds (at least until SSDs finally get around to closing the gap), and I'm kinda tempted to spring for one in my next home build, even if it'd be highly impractical for my needs (I can at least leave a slot open for one).
You could also go with newer tech for even better absolute capacity and performance (but the cost of drives hampers the economics of it a bit, unfortunately; by my math, $3952.52 for the cheapest LTO 7 drive on Newegg + $61.30 per 6TB cartridge would push that crossover point to 264TB v. retail Reds or 504TB v. shucked - that'd be a lot of data to hoard for LTO 7 to be worth it unless you can find a screamin' deal on a drive on eBay, though I wouldn't wanna shuck 63 drives, either, lol). Going the other direction, LTO 2 has lower up-front costs, but the media kills any economic benefit (worse price/GB than hard drives). LTO 5 seems like the sweet spot there, for now at least.
Of course, this is going a fair bit beyond "consumer grade", but I do know of quite a few serious data hoarders who've stuck with tape for this exact reason.