Western Digital admits 2TB-6TB WD Red NAS drives use shingled magnetic recording
blocksandfiles.com
blocksandfiles.com
Only up until the array needs to be rebuilt (eg replacing a drive). At which point the workload is the literal opposite of "bursty with idle time". :(
My understanding is that it removes the ability to do random writes, since you'll disturb the nearby tracks. Like drawing pen thickness rings with a sharpie by overlapping them.
Instead you read data in, modify it, then write it out, like an extreme version of updating a byte in a cache line.
Your ability to read is more precise, so you have no trouble reading data laid out in this way.
Large contiguous writes are also pretty quick. I've owned one, and if you stick to large files it's fine. Pathological workloads like creating millions of small files can take days.
Typically stuff you write randomly (small files, databases) one tends to read randomly, so it's usually worth linearizing writes.
The only case where you are at a performance disadvantage is when you did random writes followed by more than one linear read.
Things like LogFS do this, as do most SSD's. I'm pretty surprised modern HD's don't do it internally too.
Are there filesystems or databases which are specially designed to optimize for this constraint? It would basically boil down to structuring your whole system as a set of interlinked mostly-append journals.
Reads of a single record usually only read a single disk block. In-memory indexes to locate the block are fairly small, even for 1B+ records.
Reads of a range of records do N range reads of blocks, where N is typically < 5, and the number of bytes read compared to the number of bytes returned to the user is typically <10% overhead.
In fact OSes have been doing this for years now. Over a decade. There is a drawback however. To make this work you need to buffer your writes in a cache, typically in volatile memory, so they can be sorted before being written to the drive. The larger the cache the more improvement you see (although there are diminishing returns on this), but also the more data you lose in the event of power loss or kernel crash. So it is a tradeoff between speed and resilience.
Interesting thought about how this changes for SSDs. While you don't have seek times to consider anymore, if you have two writes that happen to fall on the same block in the SSD it is a big win to combine them, so SSDs still want to buffer data and have to make the speed/resilience tradeoff.
Then keep some accounting information elsewhere to map the real location of data to the place it would have been written had linearization not been done.
Still, I did recently reshape from 4x6TB raid5 to 5x6TB raid6. The new drive was SMR WD Red, it took 12 days and I was quite pissed when I found out the SMR bit. It wasn't in the datasheet, if I knew upfront, I would never bought it.
SMR has a terrible rep for a reason, and no one with any knowlege of what SMR is would buy which is why they have to be unethical by removing this spec from the spec sheets
SMR is cheaper for them to manufacture, and is TERRIBLE technology that really has very little actual use case, except Archive storage. The manufacturers want to push SMR for more than that because it is more profitable for them
It's not that terrible if the customer knows what they're getting. It's just that SMR competes with tape more so than with conventional (CMR) hard drives. Many workloads are good enough for SMR.
One workload is, which is what you said.. Archiving, what one would normally do to Tape.
Everything else is is substandard and should not be used
SMR sounds like it's great (value) at providing random read access to large amounts of data if you only have sequential writes or infrequent writes.
Things like data warehousing, where you would want to store on something else for the current time period, but could store a big blob as the data gets finalized (once all nodes have reported in, or maybe after N days if you drop some columns at that time). Or household archival where data basically doesn't get changed once it's put on the storage --- wait for a block sized write to accumulate and then do it (need a different tier for the intermediate writes... Some SMR drives have a PMR section for this too).
But, device managed SMR is not a very good solution, even if the use case would be good for SMR, because the filesystem is most likely not aware of the need to cluster writes.
I mean, it could pull-off global block remapping, I just didn't expect it.
During that time, data is not contiguous. Nor is there any guarantee that it will ever be.
SMR is way more like tape than an SSD.
Also tape drives are too expensive for normal consumers.
Between cigarettes, credit default swaps, and asbestos, I think this is a phrase we should become very wary of saying with a straight face.
I have a couple of 8 TB Seagate SMR drives i use for backup, and for that purpose they've been great. They're cheaper than PMR drives, and i honestly don't care if the backup takes an hour extra to complete.
Unless you're in a business scenario, how much of your data on your x*8TB drives at home do you actually write to on a regular basis ?
Half of mine is family photos, backups of clients, movie/music media, etc. For this purpose SMR also works well.
I'm far more worried about power consumption for NAS drives than write performance.
If you don't then the tape drive write speed falls off pretty quick as it does stuff to cope.
(which is often argued as something positive but I would argue is a bad aspect of ZFS (regardless of whether SMRs are used or not))
See for example "New prefetcher for sequential scrub" and "Improving resilver: results & operational impacts" from http://www.open-zfs.org/wiki/OpenZFS_Developer_Summit_2017
This is for a media server so figured SMR would be fine, but the resilver took 28hours!
No checksum errors though, so now that it's done things seem to be humming along fine.
With extra inefficiencies, yeah, SMR drives will only be a few times slower than regular drives for sustained sequential write workloads, as opposed to thousands of times slower for random writes.
Real host-managed SMR with SMR aware filesystems and storage engines might be OK, but that's not a thing in the consumer space.
I fail to see how it makes the drives any less useful only because it is revealed that it is SMR and not PMR.
It is models from 2TB to 6TB, meaning pretty much all NAS drives purchased in the past decade. I don't know about yours, by my drives perform just as well as they did yesterday.
Yes, as plainly stated in the article
> much all NAS drives purchased in the past decade
These are new models that have been introduced in the past year. You may not (probably don't) have any of them.
1. SMR didn’t really exist in the market a decade ago, definitely not for non-enterprise grade NAS drives
2. This is a fairly recent change to use DM-SMR and not explicitly market it as such
3. Your drive performance needs != to everyone else buying NAS drives needs.
As I clearly stated: my concern is about brand damage.
>However, SMR drives are not intended for random write IO use cases because the write performance is much slower than with a non-SMR drive. Therefore they are not recommended for NAS use cases featuring significant random write workloads.
[...]
>We brought all these points to Western Digital’s attention and a spokesperson told us:
[...]
>"In device-managed SMR HDDs, the drive does its internal data management during idle times. In a typical small business/home NAS environment, workloads tend to be bursty in nature, leaving sufficient idle time for garbage collection and other maintenance operations."
That seems like a pretty notable difference to me. You probably want to know that ahead of buying the drive.
Emphasis mine: a NAS featuring "significant random write"s, is not the usual NAS, and might not actually be best described as a NAS at all.
In WD's terms, they'd tell you that if you've got e.g. your /var partition mounted over SMB to your NAS, then your NAS is receiving a Desktop workload, and so you should use their Desktop drives in it, not NAS drives. (Consider: the drives backing Amazon EBS—which receive Server workloads—are certainly not "archival media" disks, despite being nominally "network-attached storage" disks.)
The use-case for WD Red (NAS) drives is online archival media storage. Like S3 without the high availability. In other words, the thing most people buy a NAS for: backups and home Plex hosting.
This is why, I assume, WD don't literally call these drives "WD NAS", but instead just call them "WD Red." Just because you put them in a NAS, doesn't mean they'll function well there. You have to be using them for an idiomatic NAS workload. (Which also means that WD Red drives fare just fine in a desktop or server, too, if you're using them there purely for an idiomatic NAS workload.)
But you've got to further subdivide the use-case: a file server for a business is still not necessarily doing "random writes." You can create and replace as many small files on a filesystem as you like, and in a modern filesystem, you'll only be doing large batched writes to the disks themselves. It's only when you're writing within existing files—i.e. writing to a database of some kind—that you see a high number of random-write IOPS hitting the disk.
Mind you, there are a lot of things that turn out to have databases in them. A Chrome profile has a couple of SQLite databases in it. An iTunes library contains a database. Etc. These are the things that—if you stuck them on your NAS, and then pointed the relevant program at them from your PC—would constitute the NAS performing a "Desktop" workload. And, obviously, if you run Postgres on your NAS—or on a system that mounts a NAS disk over iSCSI—then that disk is doing a "Server" workload.
But the average business's use of e.g. shared Excel worksheets, does not "heavy random writes" make. Document/productivity software is almost always of the "on save, do a streaming overwrite of the whole file on the remote end" variety—which, again, is not a random-write workload, since the ftruncate(2) at the beginning allows the filesystem to grab a new free extent for the overwrite to write into.
Most software that knows that it can be used in a shared-remote-disk setting is built to accommodate that shared-remote-disk setting by attempting to minimize random writes (since this kind of workload has always been slow on the kind of storage used in NASes and SANs, especially when your filesystem or software RAID does checksumming.) These days, it's only the software that absolutely can't help it—that needs random writes to be performant at all—that still does them. And it's pretty obvious, to anyone using such software, that they're not using "NAS friendly" software. Mostly because it's so dang slow!
That's still obfuscation bordering on scam IMO, if the drives have technical limitations they should be spelled out, not hidden behind broad terms like "NAS drive", especially when they're branded the same way as older drives that didn't have these limitations.
Also, you'll have to explain to me how having these drives perform correctly while rebuilding a RAID array is not a "usual NAS" scenario. Because reading TFA it seems like getting these drives in a NAS to begin with is a challenge on its own.
The clerk at the computer shop tried to upsell me on the NAS drives, chortled when I said "the "I" in RAID stands for "Inexpensive" and told me that I would regret my purchase.
It seems to me that my cheapness inadvertently paid off?
But during a re-build there is no idle time...
Yes, they do not perform anywhere like PMR on writing.
It's so bad they simply should _never_ be used on a RAID or Zpool, where they have such pathological behaviour on resilvers that can easily become the trigger for data loss.
I would never buy one of these voluntarily, as they'd cause me headaches.
So, definitely, I need to know whether what I am buying is SMR or PMR. And wd's behaviour is unacceptable.
My WD Reds do too, but that's because mine are two years old. The switch to SMR apparently only happened in the most recent model (WDx0EFAX).
I don't get that at all.
I've been in the WD camp since the 90s but Backblaze's data has always ranked them around "better than the worst" and slowly eroding my confidence in them.
For me, Backblaze's Stats have changed my perception of HGST who I've held in low regard since they were IBM with the whole Deskstar issue.
The 60Gb ones that everybody screamed about worked perfectly if you only partitioned them out to 58Gb.
Best value for money on the market at the time so long as you remembered you had a 58Gb hard drive and behaved accordingly.
The ones I got only failed on the outer edge.
Worked out nicely cost-effectiveness wise if you didn't trigger the failure mode.
What Backblaze and experience show is how drives of a certain brand performed a couple of years ago, not how the drive you just bought will perform.
IBM used to be very good until Deathstars, and now HGST, who took over is among the most reliable. I used to be very satisfied with Seagate, then ST3000DM001 happened, and now, it looks like they are slowly getting back on track. Maybe it is WD turn right now. I don't think there is a way to know until it is too late.
Personally, I like to make RAID mirrors with drives of different brands, or at least different models. Even it it is not ideal for performance, it helps making sure that not all drives fail at the same time.
The smartmontools ticket linked in the article says that Seagate is doing the same thing. Best to avoid.
> Some Seagate Barracuda Compute and Desktop disk drives use shingled magnetic recording (SMR) technology which can exhibit slow data write speeds. But Seagate documentation does not spell this out.
---
[0]: https://blocksandfiles.com/2020/04/15/seagate-2-4-and-8tb-ba...
It looks like the only surviving product names are Ultrastar and CinemaStar. Nothing is labeled HGST anymore. Most of the HGST Helium disks were SMR, if I'm not mistaken. As of now, given the two HelioSeal series, one of the two is SMR.
Side Note: Hitachi had such a great thing going with their Schoolhouse Rock meets Dino DNA commercial. I wish this would have continued. https://www.youtube.com/watch?v=xb_PyKuI7II
i do that too, and effectivele that means i had to get one of each brand available. the choices are really limited here.
I'm so sick of HDDs, I can't wait for SSD NAS to close the gap. An NVMe M2-blade NAS would be incredible!
At work, we had literal hundreds of machines failing the same week because of sudden ssd death.
Ssds are fine for speed, but for storage... maybe one day.
https://www.pcworld.com/article/2925173/debunked-your-ssd-wo...
FWIW however I bought some HGSTs a couple of years ago that are by far the loudest drives I've ever used. They've been reliable thus far, but the noise level is so bad that I had to move them away from where the people are working.
It literally could hold all of the popular software of the time, around 50 applications.
I also had one of the first CD-ROM drives at work (the Japanese director had connections at Sony.)
We were like, "How can anybody possibly ever fill 650 MB of data storage?"
We ended up mastering our own satellite data CD-ROMs to demonstrate to people what was possible.
Pictures and movies have always been the easy answer for how it's possible to fill up huge amounts of storage.
192 Power-Off_Retract_Count 0x0032 100 100 000 Old_age Always - 2502492302
193 Load_Cycle_Count 0x0012 086 086 000 Old_age Always - 142654
^^^ those are a working, happy disk; all due to a "feature" that was on in Ubuntu by default that wanted to idle/spin down the disk every other second and it took me a while to notice.On contrary I'd worn a 2TB 2.5" Seagate out in 2 years. Bought it new, put it in the same machine the 1TB HGST had been running for 3 years already, the Seagate died earlier - or at least that's what ZFS told me. It's spinning, but it throws weird errors from time to time.
I never had experience with 3.5" Hitachis.
Eons ago there were some interesting disks in the market though; if I remember correctly my father had some Fujitsu SCSI160 or 320 drives which had their normal working temperature around 50-60C.
One of the dangers of blending drives is that there are sometimes drive geometry incompatibilities that can screw up RAID. Saw someone show they had drives that are a few cylinders off and can't replace drives now with certain replacements.
Are there too many negatives there? It seems that you first say that you distrust ("no reason not to distrust") Toshiba drives, but then you say that your experience with them has involved no problems.
This is exactly my approach lately, and it works great for RAID10 arrays (where you only need 2 or 3 different vendors - one per side of the mirror - to mitigate shared bathtub curves). I usually go with SSDs nowadays, though (half Crucial, half Kingston), both for performance reasons and because they rebuild faster (and also because the limited lifetime of flash memory makes some sort of RAID essential for longevity, though this has been less of a problem with modern flash tech); if I really needed the capacity of magnetic media I'd probably opt for a WD / Toshiba split.
I’m familiar with Crucial BX500 and Kingston UV500. They are both SATA and are decent performance considering the limitations of SATA 3. If price isn’t a factor or performance is important I’d go with a Samsung 860 EVO or Pro on SATA. I’ve also used a smaller number of SanDisk by WD SATA SSDs and they work fine as well. Had a bad Samsung 840 series but I think failure to implement proper overprovisioning was to blame. Also had a bad Intel SATA SSD but it was an older generation product and had a small sample size of only a few Intel SSDs, so not writing Intel off; yet it’s suffice to say their products don’t seem price competitive with Samsung’s offerings.
As for NVME SSDs I have used and installed probably ten so far but I think every single one so far has been a Samsung 9xx m.2 and I’ve been extremely satisfied with all of them, especially the 960 series EVO and Pro and moreso the current gen 970 EVO Plus and Pro.
Both the desktop on which I'm typing this comment (my personal "gaming" rig) and my work laptop (a Thinkpad T470) are using Crucial MX500s (1TB, M.2); the former has a couple extra M.2 slots, so I'm considering migrating to a Crucial P1 + Kingston A2000 combo (thus biting the bullet and switching to NVMe, and getting RAID going on this machine). Most of my 2.5" purchases tend to be Kingston V300s unless the MX500 happens to be cheaper (e.g. most recently for an old XP computer I'm rebuilding for a family friend) or I'm specifically buying both for the aforementioned purpose of RAID diversification (which is a bit tricky, since Crucial and Kingston use different increments for capacity, but I'm also perfectly willing to pull a bit of a SSD "no-no" and use the leftover space on whichever drive for a swap partition, or else just let the space sit unused, or perhaps use it for temporary files where I don't care about any sort of data preservation).
In any case, no complaints with any of 'em. I haven't tried the P1 or A2000 yet (let alone Kingston's pricier KC2000), so I can't attest to those, but the MX500 and V300 have both proven to be reliable workhorses.
I'll definitely be sure to try Samsung again at some point; the first SSD I ever bought was a Samsung (at Fry's), and I had pretty awful experiences with it, but that was back when SSDs were a pretty new concept so I'm guessing things have generally improved.
WD Drives were actually HGST drives with a different label for several years.
I still can't fathom how companies can be so oblivious about their actual customer market. People speccing individual hard drives are integrating their own systems and want to care about the details, not just some colorful indicator of "good/better/best".
They likely rebate / warranty their important hyperscale cloud, enterprise, & OEM customers. According to whatever custom contract they've negotiated.
And don't really care about the retail segment, given the margins. Most people probably won't notice. And those that do won't be able to do anything.
And yeah individuals can never really do anything in the short term. Eventually there will be a suite of best practices to test a new drive to make sure it's not SMR, and likely some tweaks to filesystems to ratelimit rebuilds, based on knowing which drives are actually SMR. But they're arbitraging away brand loyalty to WD Red, for a short gain. Given the goodwill they enjoy(ed) from that Blackblaze study creating a refrain of "don't buy Seagate", this just seems foolish.
If risk and failure were spread more evenly, systems would behave more like gas and less like a ceramic.
A typical size for a SMR zone is 256MB, and any writes that aren't appending sequentially to what's already been written to that zone count as "random writes" on a SMR drive. That's an absolutely massive block size to expect filesystems to work around, and I doubt there are any commonly used filesystems that can store their metadata in a SMR-friendly way.
SMR is extremely painful for anyone except those that can handle distributed storage systems. I think the only place where SMR makes sense for home users is in a DVR, or game assets where a virtual disk is streamed down from a cloud provider in a single pass, and no interior mutations.
HN is fringe, I am not making statements about people with drobos.
Shucking has always had weird economics. I think mostly because it's used by the manufacturers as a release-value for excess capacity, in an attempt to somewhat avoid the memory chip boom-bust cycle.
Now I am having 3xToshiba drives and one Hitachi and for now I am happy.
Would a class action suit help to prevent this kind of action in the future? Perhaps the punitive fine will make the cost calculations that led to this decision different since they will need to include the risk of judgement in the costs.
Reviews these days are rather all about GPU's and CPU's.
And that work needs to be paid for.
And nobody buys review magazines any more.
And advertising revenue on review sites can only pay for a small handful of SKUs to be reviewed. So only the sexiest, mass-marketest units get reviewed.
And the vendors fly under the review radar by flooding the market with so many variants that there isn’t enough funding to review more than a tiny fraction of them.
And by the time any headway is made into reviewing current models, new ones are released and suddenly nobody is interested in the old ones, further squeezing the window of opportunity for advertising revenue. And eliminating entirely the incentive for long-term testing.
Meanwhile, review sites realise they can get paid more in bribes/freebies/access than in advertising, and either outright compromise their reviews or only report the hits and never the misses.
And so consumers are flying blind, trying to extract signal from the noise of furious churn.
Pretty much. Like the rest of news, accuracy is expensive, entertainment is cheaper, and anyway the audience prefer the latter.
I think this is true for almost any product nowadays that isn't in a ridiculously small niche.
Most non-niche products on Amazon I've found get a poor score on "FakeSpot".
I believe someone even admitted on reddit that he/she was actively creating fake reviews and fake stories about certain boot brands on relevant subreddits.
Unfortunately you can't trust anything online anymore. You need to spend hours or days researching the technology or product category in order to make a judgement on the products available.
If you want to get really thorough, automation starts to be infeasible and unreliable: http://blog.stuffedcow.net/2019/09/hard-disk-geometry-microb...
At least in IT we have Stallman-esque zealots who refuse to be bribed. I don't think we have the equivalent in other markets unfortunately.
For performance I think the problem is that not every single capacity is reviewed. Typically the top capacity in each range.
[1] https://docs.microsoft.com/en-us/azure/storage/common/storag...
RAID is only effective for workloads small enough that you care about a single machine.
https://www.backblaze.com/b2/storage-pod.html
"For Backblaze Vault Storage Pods each is one of 20 pods needed to create a Backblaze Vault. A Backblaze Vault divides up a file into 20 pieces (17 data and 3 parity) and places a piece of the file on each of the 20 Storage Pods in the Vault. We use our own implementation of Reed-Solomon to encode and distribute the files across the 20 pods, achieving 99.99999% data durability. We open-sourced our Reed-Solomon encoding implementation as well."
Maybe they use RAID for the OS.
Backblaze's main reliability system is erasure coding[4] of shards of files, which is not RAID[1][2]. RAID mirrors disks or volumes on block or filesystem layers, not files in userspace.
That being said, I stand corrected in that they did use RAID at some point in time[3] (RAID6 on 13+2 drives), and may still be using some form of RAID. OTOH their reliability calculations don't seem to consider RAID - calculating with individual drive failures and shard rebuilds.
[1] https://en.wikipedia.org/wiki/RAID
[2] https://en.wikipedia.org/wiki/Non-RAID_drive_architectures
[3] https://www.backblaze.com/b2/storage-pod.html
[4] https://www.backblaze.com/blog/cloud-storage-durability/
Slow performance I could understand (though RAID rebuild is usually sequential, not random), but how does SMR cause these "dropouts"?
The same happened with WD Green drives, which are not graded for RAID. Their error correction logic typically allows for a lot more attempts to read the data than a datacentre drive, resulting in timeouts and the drive dropping out of the array (which is why it is a very bad idea to put a consumer drive in a hardware RAID array).
Now WD Red NAS are meant for RAID arrays but I suspect WD assumed they would be used for software RAID only, which typically doesn't have those timeouts. But if so, it should be clearly stated.
The typical bad scenario (and I have been burned by this) is that let's say you use RAID5, and one sector goes bad. While the disk tries to read it, the RAID controller kicks the disk out of the array because of the timeout. Now you need to replace the disk and rebuild the array. In the rebuilt, as you are doing a full read or all disks, you are pretty likely to find another sector that is slow to read on one of the other disks (particularly on multi-TB disks). And then the controller will kick another disk from the array. And now you lost everything.
Also why data scrubbing is pretty important in NAS.
https://raid.wiki.kernel.org/index.php/Timeout_Mismatch
It's important to know this kernel command timer and SCT ERC mismatch is (a) common and (b) affects mdadm, LVM, Btrfs and maybe ZFS RAID on Linux. I'm not really sure whether the kernel command timer applies to ZoL or if a vdev has its own policy.
Some people might be able to put these drives to good use, but only if they know up front that it's this kind of drive they are getting.
Why does that happen?
Because with DM-SMR drives there is usually a non-shingled area on the drive used as a buffer. Once the buffer is full, performance tanks dramatically as tracks are being rewritten. This is especially true in my experience when writing lots of smallish files vs several very large files.
https://www.guru3d.com/news-story/cpc-hardware-amd-ryzen-7-2...
Manufacturers have slowly started realizing this circumstance and are obviously able to exploit it by making products cheaper to manufacture down the line.
I believe this is a known technique in mass-produced markets. Called 'debasing' or something like that. Even new cars have more bells and whistles when first released while as the model gets a stead consumer base the manufacturer might reduce accessories or get cheaper versions, etc.
If you pay too much for too little on your car, oh well. Its probably too late to do anything about that now or you would have returned it to the seller already. For most people they settle for enough car for enough money. You can always add aftermarket parts or customize it, because there is a large secondary market waiting to make your outside investment even more worthwhile to your own judgement. If you pay too much for soap, it’s not even worth it to return it to the store in your time or your money so the market simply wins. It’s a rigged game akin to Las Vegas house games.
As buyers it’s just stressful being so wary all the time. I don’t know what the better way is but surely it must exist, we just haven’t made it profitable enough yet for the right people.
Most of the updates like this in the SSD market are actually pretty harmless—switching a SATA drive from 64L to 96L TLC generally isn't going to change the performance or power characteristics enough to care about, and is more likely to be beneficial than harmful. But when they start introducing QLC NAND into a product line that originally was TLC-only, that's a problem.
Were these new enough that anyone bothered?
At the moment everything is sitting behind the bogo SATA disk interface which was never designed for this type of storage - rapid burst writes to the 20GB onboard buffer area and than very slow I/O as it is emptied out.
Surely a native OS interface would be better for these drives.
Some file systems (notably ZFS and Btrfs) use a copy-on-write approach that should in theory not require many random writes, but I don't think any of them are adapted for the large blocks you need to write to for SMR. (256MB?)
It should be possible to make a filesystem with drastically better performance on SMR drives given the right low-level access, probably based on something more like a log-structured merge-tree, like e.g. LevelDB.
One interesting example I found was Dropbox which seems to store data on SMR in big immutable chunks using raw Zoned Block Command access. https://dropbox.tech/infrastructure/smr-what-we-learned-in-o...
The filesystem would need to do periodic compaction. You would necessarily encounter situations where a write to “allocated” data would fail because the disk is out of space, which is a new, surprising error that would undoubtedly confuse some user-space programs.
All file systems take up a bit of space to store their own metadata etc. So users are already used to 'lose' a bit of space on formatting.
The only solutions I see here are to block IO until the compaction algorithm frees space, or return ENOSPC. Both options violate assumptions that real-world programs make.
They help with poor random write performance only at the cost of poor sequential read performance.
Both SMR drives and SSDs benefit greatly from sequential reads. A file that was randomly written will be read back in random order regardless of the reader's interest.
Sounds like a perfect use-case for log-based filesystems.
But for SMR, the reading track is narrower than writing track. For example, reading track's width is 1 unit, writing track's width is 2 unit, there is a 1 unit overlap between two adjacent tracks.
And here is how it works, assuming writing sequentially: 1. Start with blank disk, and write to track #1. This is happy case, as no data there, we can simply write to it. 2. Write to track #2. Still can simply write to it. However, this will overwrite lower half of track #1, but as long as reading only need 1 unit, track #1 record is still good. 3. Write to the following tracks one by one, all good, same story with step #2, over writing half of previous track. 4. Repeat until track #n. 5. Okay, all of a sudden, now we need to change something in track #1. What would happen now: we cannot simply overwrite the stuff in track #1 as this will overwrite the top half of track #2, which is the reading track of track #2. So we'll have to read out track #2 and and write track #1 and then write back track #2, which in turn has to read out track #3 and write track #2 and write back track #3... 6. Repeat the previous operation until all impacted tracks are written back.
In reality, the situation could be slightly better: if the disk is almost empty, the firmware can always find somewhere to write to without reading out and writing back. But if the disk is relatively full, then the case could be much worse: every writing could technically trigger another reading out and writing back operation. So how long would it take to finish a writing operation can hardly be predicted.
The worse case is that the disk is full and the writing is targeting whole track #1, then almost the entire disk need to be refreshed so that it can complete the operation...
Thanks for explaining.
It would be amazing for someone to come up with a “worst case write” benchmark for SMR drives.
[0] http://blog.schmorp.de/data/smr/fast15-paper-aghayev.pdf
The band size of 17-30MiB is a lot less than the 256MiB mentioned here by others.
I have them in a 4U chassis in the attic, with 5 case fans blowing through; two on the rear and three across all the bays. It's a 24 bay box, but I've not yet had pressing need to fill the rest of the bays.
At the time, WD was assimilating HGST's drive tech where there were cost/performance/reliability benefits to the point where quite a few 'WD' products were just relabeled HGST ones, possibly with minor firmware/logic board tweaks. I seem to recall that at the time some WD Black, and maybe Gold, models at least were rebadged HGST units. The aim at the time was to fade out the HGST brand so that everything was WD and I pretty much think that has since been done.
It sucks to have to qualify drives (again?) in small businesses, thought this was a solved problem. It's more offensive they're doing it with the smaller capacities, too.
(I should note SSDs don't stand up to manufacturer's claims in error scenarios, too. But that's fairly well known I think?)
[0] https://blocksandfiles.com/2020/04/15/seagate-2-4-and-8tb-ba...
- https://www.backblaze.com/blog/hard-drive-stats-q2-2019/
- https://blocksandfiles.com/2020/02/13/seagate-12tb-drive-fai...
What matters in that context is something similar to $/GB x (% of drives which make it to end of useful life in their application context). Cheaper drives which are less reliable might well win on that score. They might also expect to retire drives rather sooner than an average NAS user as the power cost of running older, smaller drives becomes an issue.
In an average small NAS setup the cost of a failed drive is potentially much higher. You incur at a minimum a performance penalty while the array is degraded and a significant performance penalty while it's rebuilding and have lower guarantees against data loss during that period and you can't average that out over thousands of drives.
Do you realize how many different models there are, customized for different workloads? Do you realize how many different sets of firmware there are?
It’s like saying Honda has shit cars because one model year had a problem.
Numerous new technologies led to severe issues like : 5400.6/7200.4 platter dust issues killing heads the more they heat, the infamous LBA 0 fw issue on desktop drives, motor bearing seizure on 7200.10/12, lots of bad batchs on 7200.14 ST DM001/002/003, more than WD or even Hitachi. On SAS drives, the Cheetah 15K.x are not very reliable (replaced LOTS of ST3600057SS), but the Fujitsu and Hitachi are not very either.
Toshiba drives are not great either. The 2.5" MK series had an important return rate, more than the competitors. The MQ01ABD is better, but faces platters demagnetization that leads to bad sectors. The 3.5" ACA drives are a Hitachi prod line that was "given" to Toshiba when WD acquired HGST. They are quite basically Hitachi drives with a Toshiba firmware.
Edit : and now DR on Seagate drives like the Rosewood family (like the ST1000LM035) is the absolute worst nighmare for any data recovery guy. Ultra brittle head stack, easy platter damage when shocked, self encrypted firmware and SMR. If those disks makes some motor noise or head clicking, we don't even bother to open them anymore.
Customers here often mistake redondancy for backups. Had to save a 8 years-work worth excel file on a failed Kingston USB key today.
[1] These were bought from 3 different retailers at slightly different times, so it's not a "bad batch" either.
Anecdotally, the only Seagate disks that have caused me significant headache are the Archive line that definitely uses SMR and is marketed appropriately because of it.
Compared to their high end consumer IronWolf Pro/NAS series they have better tech, no RAID drive limit, higher rated workload, and longer estimated MTBF. Looks even better now with the 14TB Ironwolf at $460 and the 14TB Exos at $340 on Amazon.
This does not make sense - why smaller disks are SMR, while larger ones are CMR? Perhaps they just switched that in comment?
Maybe the market for 8-14TB is different enough that it warrants CMS? For instance maybe the 2-6TB is for hobbyists/small companies who might not realize the drives behave differently while the larger and more expensive drives are more often use in large companies with dedicated IT who would realize something is off?
Given that TFA mainly mentions 2-6TB several times I doubt they got it wrong.
or that the 8-14 where used in "serious business" where you could not fake your way around and the smaller ones for consumers or smaller firms would not notice or could not cause bigger harm.
SMR is going to reduce platter size and quality requirements for a given size in GB (because the data is packed more densely). A given capacity can require fewer platters/sides, or a blemished platter can now be used where it couldn’t previously.
No, I think that's correct. It's definitely weird, but it's consistent with what users have observed.
That way they can come out with a more expensive PMR SKU in the NAS series. SSD manufacturers did a lot of the same where features were stripped from consumer drives (PLP, DRAT/DZAT, etc.), segmented into a new market (notice the NAS SSDs?), and sold back to us as more expensive SKU.
[0]: https://documents.westerndigital.com/content/dam/doc-library...
It's so frustrating that all that was needed to avoid most of this was some transparency. At least list in the specs that they have SMR. It's essentially lying by omission.
But I'd probably go a step further and say that marketing an SMR drive as a NAS drive is user-hostile and destined to cause problems
Why DMSMR became a thing in the first place? Windows! And non-upgradeable proprietary storage array hardware.
It seems that business people in those storage companies have put big money into SMR, but only realised that an extremely big chunk of the market cannot adopt SMR after drives shipped. And so they simply slapped a cheap and easy firmware hack on top of SMR drives to make them work with non-SMR aware software.
Just supporting TRIM and exposing/documenting the SMR layout/regions would be enough, I think. I'd like to buy that.
From salesmen of "enterprise class" overpriced hardware.
https://www.westerndigital.com/products/data-center-drives/u...
As far as I can tell it's seemingly impossible for an individual to buy host-managed SMR (ZBC) drives. Presumably they're only sold in bulk to datacentres. Maybe that could change if some demand materialised and was communicated to a suitable retailer...
As I alluded to above, HDDs aren't the only things affected by this trend; SSDs, or more precisely NAND flash manufacturers, have become equally secretive about endurance and data retention. Older datasheets were freely available and specified the number of cycles each block could be guaranteed to be rewritten and how long the data would stay, but datasheets for newer flash are basically all NDA-only and even then you won't get the exact details. At least in that case, I suspect they are trying to hide the inconvenient truth of decreasing reliability; software workarounds can only go so far.
It feels like the first two are increasingly taken, so businesses are lazily switching to various versions of #3. I don't know where the regulators and punishment for doing that have gone.
I use WD Reds in my Synology NAS in a SHR configuration for daily backups. Should I be worried? Should I replace the disks with other models? What are the risks if I keep the current configuration?
However, they have cut corners / optimized costs that assumes the drive is not under constant pressure, but rather mostly idle. What this implies is that they likely just have a hybrid magnetic/ssd backend, which is what the outrage is about.
No need for hybrids. DM-SMR drives usually have a non-SMR zone acting as a write buffer, similar to the SLC write buffers in non-SLC SSDs.
With my limited understanding, the new models will fail under intensive work-loads. For example adding them to existing RAID setups to replace a failed drive, where in the beginning they have to be sync-ed and they fail under load after a few hours (either by dropping in performance or not being recognized at all).
They will work reasonable well under lower loads (more common for home NAS setups).
The biggest problem, and what the article is trying to point out, is that Western Digital (and Seagate) are doing all they can to hide this info from customers. They even advertise the HDD as being NAS/RAID friendly, when they are clearly not meant for that type of loads.
Unfortunately, the nature of RAID means that you can be dealing with a lot of random writes. Further, when a drive fails and you want to replace it, you are pushing massive amounts of writes to the new disk.
A lot of this is due to the Shingling. Like shingles on your roof, the physical locations where data is written overlap. The problem is that if you need to update an already written location, you may also need to update surrounding bits. This is where the random write performance degradation comes from.
https://blocksandfiles.com/2020/04/15/shingled-drives-have-n...
Is it so they can use fewer platters to get the same capacity, thus reducing the manufacturing cost of the drive? Or is there some other technical reason to do it on low capacity drives?
"However, SMR drives are not intended for random write IO use cases because the write performance is much slower than with a non-SMR drive."
Might this be a way to detect SMR usage by allegedly non-SMR drives?
Do a whole lot of random write IO -- and time and statistically sort the results?
My assumption would be that the Purple drives become the default for NAS usage as the surveillance use-case much more closely matches the requirements of a RAID rebuild (a fairly constant write-stream).
My anecdata sample size of 3 has WD Red drives priced very closely to equivalently sized WD Purple drives, with WD Black drives a significant percentage more expensive (40% +).
Does this mean WD Black drives are the last hold-out against SMR for WD brand?
I hope they have a class action suit coming.
I just put a ST8000DM004 8TB drive into my 12 disk raidz2, on zfs 0.8, it took 8 days (so terribly slow 9MB/sec), but no errors resilvering
A ST8000DM004 is capable of linear reads above 200MB/s, and a normal drive should be able to handle the same speeds with linear writes.
I would not recommend those drives for use with ZFS from my own experience: one of my drives disappeared after 1.5TB of writes while moving a dataset, and only went online again (with a SMART failure) after a power cycle.
Up to now I was assuming any random 7200 rpm drive will do. Apparently I now need to be extremely cautious?
Their larger disks are still claimed to be CMR at the moment, though how much credibility anyone should give a WD spokesperson at this point is obviously open to debate.
* 256 MB (or more) - most likely it’s an SMR drive
* 64-128 MB almost guaranteed it’s an PMR drive
I base this on my observation, that I’ve never seen an SMR drive with less than 256 MB of cache.
"Designed to handle workloads up to 550TB per year, the Ultrastar DC HC520 is the industry’s first 12TB drive and uses traditional perpendicular magnetic recording (PMR) technology to make it dropin ready for any enterprise-capacity application or environment."
[0]: https://documents.westerndigital.com/content/dam/doc-library... (search PMR)
Luckily that problem only led to the largest (so far!) recession / crisis humans have seen.
Ok snark over, but we have seen things like this before - take steam boiler production mid to late 19C. This took insurance companies simply refusing to insure anyone who did not manufacture boilers to their standards to clean up.
Something similar is needed now. Do you know why you cannot claim your covid-19 cancelled holiday from your insurance - because pandemics are not covered by travel insurance - do you know why? Because Re-insurance companies look at the likelihood of pandemics and simply refuse.
I think something different is needed - I think insurance is related to it. Basically we treat government as the insurer of last resort - but we have no real idea of the costs we are shouldering - or the risks we are ignoring.
So what if we had to have obligatory true cost insurance? Buy a house? great is it insured against flooding? Now you built on a flood plain like half of England. Fine your premiums are now twice the cost of your mortgage.
Are you selling crappy products through your business? Great please show business insurance that will make the people you rip off whole. Don't have that insurance? good luck getting anyone to buy off you.
Ok ok it's still a idea very early alpha
I believe it's not likelihood alone. It's a combination of very low likelihood, very high cost event that doesn't make it worth it for most.
In the end it likely wasn't the re-insurers decision. They cover pretty much everything for the right price, also pandemics. But no end-customer would've paid the extra price for pandemic coverage so no insurers used it. Unless it's outright illegal I've yet to hear of an event that no insurance company would cover.
> The funding for the US $10,000,000 prize was unconventional in being "backed by an insurance policy to guarantee that the $10 million is in place on the day that the prize is won."
Lots of people seem happy enough taking the risk of slightly crappy products?
Established brands with a long running reputation is another way to achieve similar goals to your insurance proposal.
> Some users are experiencing problems adding the latest WD Red NAS drives to RAID arrays
No, these drives are not 'fine' for NAS use
>keep[s] getting kicked out of RAID arrays due to errors during resilvering
This problem is independent of what type of workload you actually use the NAS for.
This is hyperbole, but "network attached storage" could basically refer to every hard drive now and not be false advertising. No matter how unacceptable you view this, it's the kind of shit hardware manufacturer's often do.