Western Digital’s SMR disks won’t work for ZFS, but they’re okay for most NASes
arstechnica.com
arstechnica.com
> Conclusions
> We want to be very clear: we agree with Seagate's Greg Belloni, who stated on the company's behalf that they "do not recommend SMR for NAS applications." At absolute best, SMR disks underperform significantly in comparison to CMR disks; at their worst, they can fall flat on their face so badly that they may be mistakenly detected as failed hardware.
That's definitely not saying it's "okay for most NASes." Could we get a rename? It looks like Ars Technica themselves might be going back and forth; the current URL contains the slug, "western-digitals-smr-disks-arent-great-but-theyre-not-garbage," indicating that may have been the old title.
Source: it’s a well-known fact that editors have disclosed many times over. Also, as a subscriber I often see a different title in my RSS feed.
https://arstechnica.com/information-technology/2016/10/tfw-a...
To be honest performance is as I'd expect out of a hard drive. But I've mostly been using SSDs on my PCs.
I've put my instrument VSTs on it and I'm using the Windows NFS client (you use cli command to mount after enabling in add windows features)
Standard windows file sharing (SMB) seems to top out at 25MB/s and I've no idea why.
I'm not too worried about it being SMR at least looking at my use case.
But it's really messed up that WD hid the technology used.
If anything it would drive sales for their more expensive products.
Also it does highlight that in a lot of tech product fields we don't have good reviews. Because there's no money in it.
Like with laptops there's notebookcheck doing technical reviews.
With displays rtings and tftcentral.
But with other less used stuff it's a YouTube video from someone saying "yeah it's a harddrive" and reading the marketing.
Or a drive will get reviewed. Then the manufacture will start selling something with the same brand but different model number to confuse customers.
But nobody will re-review, partially because they weren't told.
I am looking to sell my DS220j however. The lack of support for btrfs is very unfortunate, as a NAS shouldn't be using ext4 with zero integrity measures and zero bit-rot protection in 2020.
RAID1 will not save anything as if there is bitrot, without checksums there is no way to know which copy is correct.
I'm looking to replace the DS220j with the DS220+ when it comes out, which is supposedly in just 1 week based on a local PC shop I called. I also might upgrade to the DS420+ depending on pricing.
Yes. But 220+ is way more expensive than 220J. For Something I expect in 2020 should be standard feature.
And those Ironwolf HDD aren't cheap either. The way things are going Cloud Storage is an increasingly attractive option.
The DS220+ does not stack up favorably on the pricing front though.
The problem occurs when you are dealing with sequential files with small sector sizes, and non-sequential files, where performance is abysmal: 93% slower than CMR as benchmarked by servethehome and Ars.
This PR looks like adding support for traditional sequential rebuild way like Linux MD that suitable for SMR disks.
Further, I'm not aware of a single raid system built in the last 20 years that doesn't do online rebuild. This means its only sequential IO as long as the RAID controller isn't servicing other IO requests. Futher it also assumes that your running a dead simple RAID system that is RAIDing the entire array rather than doing the RAID on top of a LVM/thin provisioning like setup (this is basically ZFS's problem).
Frankly, I would be willing to bet that its possible to tweak a benchmark like iometer to use a sufficiently large part of the disk with a 50% random RW workload and cause quite a number of RAID controllers to kick the drives for write timeouts. Some percentage of them might even kick multiple disks at the same time resulting in complete array loss.
Wait, what? Why? Because it's too slow?
> we can model an ideal ZFS resilvering workload with a massive sequential write — and we did exactly that, using 32KiB blocks of incompressible pseudorandom data (...) With this test workload, we achieved a throughput of 209.3MiB/sec on the Ironwolf, but only 13.2MiB/sec on the Red — a 15.9:1 slowdown
So if I understood correctly "very large block sequential write test" is only 16% slower, but reducing the size of the written block to 32KiB, like apparently ZFS does, it becomes 16 times slower?
Couldn't that be fixed in ZFS? 32 KiB block is too small for SSDs too, if I understand correctly. (Edit: I refer to: https://tytso.livejournal.com/2009/02/20/ "However, with SSD’s (...) you need to align partitions on at least 128k boundaries for maximum efficiency." -- it seems that 128K was an important number?).
If the writes are sequential, some kind of caching (i.e., blocking the sequential writes) could be made to avoid the slowdown? How big were the blocks written when the 16% slowdown happened?
And what is the minimal needed block size to avoid many-times slowdown?
sudo zfs set recordsize=[size] data/media/series
This can be done at any time, but will only apply to new files.ZFS is a Copy on Write system, which means that if you take a snapshot (or duplicate) a file and modify one byte, the modifications (min size of a record size) is created then (so 128KiB default). Each block is then distributed reduced to about ~32KiB of writes because of striping, parity, etc.
If you made the record size something large, let's say 1MB, then it would mean that changing 1 byte of a duplicated or snapshotted 1GB file would cause an additional 1MB of storage. Very, very very inefficient. However, still better than a non-CoW system, where duplicates and snapshots would occupy an additional 1GB irrespective of modifications.
So to summarize, ZFS is intentionally designed and benefits from small block sizes, and rather than changing it, you should simply buy a non-SMR drive esp. because they are available at comparable prices.
Also keep in mind that an advanced format HDD sector is 4096 bytes. I haven't heard of any drives with more than ~4096 bytes per sector. So as long as your block size is an integer multiple of 4096 bytes, you are getting all the benefits of your HDD medium.
I believe SSD native sectors are similar, so again, there's not reason to go for huge sectors. SMR's abysmal performance with small sectors is a defect, something you should RMA the moment you receive a SMR drive (WD is honoring replacements, just mention the magic word of 'ZFS rebuild' to support).
There are some NVMEs which are 8192. Not sure which, because, we haven't learned a damn thing in ~15 years.
I distinctly remember back in the 2000s Adaptec giving Theo at OpenBSD a run-around about driver code because they had lawyers in their ear trying to convince them that they should be patent and licensing trolls rather than hardware vendors (worked great for Lucent) and finally, in retaliation, he shipped a version of OpenBSD without Adaptec RAID support.
https://openzfs.org/wiki/Performance_tuning#Alignment_Shift_...
Why is this relevant? Because some of these NVME drives report wrong values for their sector size, so we're still in that same boat of being flat out lied to by the people we buy this hardware from.
Some NVME SSDs, per testing, report 4096 when they're actually 8192. So you don't know what value to use until you set up a pool/dataset, test it for performance, and then if the performance seems off versus say... Anandtech benchmarks, destroy the pool/dataset and try the other value (this becomes ashift=12 or ashift=13 currently).
You can find discussion on the ZFS subreddit about this pretty commonly, and also on the PostGIS and OpenStreetMap user groups and github issue forums since those people tend to be on the bleeding edge of the disk performance market.
SSD "erase" operations involve fairly large sectors but these are something you'd want to limit anyway, for the sake of data resilience. So using small sectors for SSD's may make sense too.
1MB recordsizes make a lot of sense for drives that primarily hold large files for archive or preservation. That lines up with the intended use of SMR drives. The problem here is that these were stealth-SMR drives, saving WD money at the customer's pain.
For the small % of use cases where you have large files that are written sequentially, and never edited internally, labeled-SMR with large records can be a good solution that saves money.
They are still hiding fundamental parameters needed to make these drives work properly (number/size of SMR regions, how much CMR is available, and ways to read the percentage the cache+CMR is full at any given time).
Perhaps a good analog for an SMR drive would be some of the early hierarchical storage servers that served data from RAM, then from spinning metal, and then from either tape or optical disks, each being slower than the previous. You could get fast writes until the RAM was full, then slower, disk-like writes while the HDs were getting filled, then excruciatingly slow throughput when you needed to hit the MO drives with little RAM or disk storage to help you.
* Append-only * Only use every other track
If used like that, they can be considered just another hard drive, with an unusually large gap between tracks.
But if you try to pack on as much data as the drive will hold, and then modify parts in the middle, the drive is going to have a bad time.
I don't like that they hid it. If they want to keep SMR around for their 2-6TB Red variants, I think they should reclassify those drives with another "color" and not certify them for server applications (WD Pink?)
[1] https://blog.westerndigital.com/wd-red-nas-drives/
[2] https://www.amazon.com/Red-4TB-Internal-Hard-Drive/dp/B07MYL...
[3] https://www.amazon.com/Red-12TB-Internal-Hard-Drive/dp/B07RQ...
What is the size of these blocks? And if the problems occur during raid rebuild, aren't the rebuilds writing data sequentially?
And, I had had failures of some of the drives before, but when I pulled them and did extensive tests of them, they were always OK.
So the drive may not have failed, the RAID just determined that the command was taking to long, fired a reset at the drive and when it didn't immediately respond marked it bad.
Since they hid the change to SMR, what else has been hidden?
Dodgy relationships between companies and journalists in tech are not unheard of, but they need evidence.
There is an argument to be made that ZFS might be getting a bit to much hype as the future of mass market storage given how dependent it seems to be on the hardware being "perfect".
Given that ZFS will happily run on pretty much anything, I think the bar for hardware dependency is slightly lower than "perfect" and somewhere around "not completely crap".