Durability: NVMe Disks
evanjones.ca
evanjones.ca
For $300, it should have been a better product. Only 6 months later the competition was getting only $100 for them.
Intel I would only use Enterprise Class gear, and even then would look to other vendors for storage
Intel trades off their name, and puts out alot of crap at the lower end which they get away with because "No one ever got fired for buying intel"
Alot of these big names do this....
Although looking at Samsung's smartphones and laptops, they might be skimping on quality, as well :/
Much as how the introduction of SMR allowed for the prices of real hard drives to increase by 25% or more, the development of new data-destroying technologies for consumer SSDs has pushed the effective price for non-garbage SSD storage up to the $0.75-$1 USD per GB level of enterprise SSDs.
Expensive enterprise SSDs make crash-proof data protection automatic. Consumer SSDs require the host system's software to explicitly flush the write cache when necessary. This tradeoff works extremely well in practice, and consumer SSDs don't have serious data loss problems for consumer workloads. Enterprise workloads that are much more paranoid about syncing every single transaction cannot safely use consumer SSDs without unacceptable performance loss, but that certainly doesn't mean that consumer SSDs are playing fast and loose with your data safety.
Although I've had problems with hardware hanging so completely that it does not respond to the reset button, and the only option is to cut power, so UPS does not provide full protection.
E.g., https://blog.tripplite.com/line-interactive-vs-on-line-ups-s...
A lot of them are pretty similar; 19V - 21V seems quite common.
Until recently phones were standardized on 5V. Now, it's a mess again. :-(
Then a battery pack could provide that, as DC.
Intel's Optane SSDs don't need power loss protection capacitors because they don't cache writes—their 3D XPoint memory is more or less fast enough to not require it. These drives correctly advertise that they do not have a volatile write cache. I have not encountered a consumer NVMe drive that lies about having a volatile write cache, but I have not attempted to test how they behave when the host requests that the write cache be disabled.
I stopped trusting Crucial when they released the P1 though. That QLC drive is terrible on every metric. Performance falls off a steep cliff once you do any writes at all, and often never quite recovers. Tail latencies are in the 1 second range, which is insane for an SSD.
https://www.intel.com/content/www/us/en/products/memory-stor...
I found this review which makes it sound pretty reasonable:
https://www.tomshardware.com/reviews/intel-h10-qlc-flash-opt...
Server grace storage will specify if power loss protection is included on the board. You can usually identify the capacitors as well. For example, they are the yellow rectangles near the edge of this Intel SSD: https://ark.intel.com/content/www/us/en/ark/products/96932/i...
Note that the Intel drive with PLP is 110mm long, which is 30mm longer than most consumer M.2 SSDs due to the additional capacitors. This won’t fit on certain consumer motherboard.
From diving into this recently in speccing out a budget system for myself, it looks to be mostly the lower end budget chipsets and motherboards that don't support it. At least that's what I've seen of the regular sized ATX boards, but it may be a different story for Micro-ATX or Mini-ITX.
The relevant info is always in the specs, and will look something like:
M.2 Slots
2242/2260/2280/22110 M-key
2242/2260/2280 M-key
For each of those, the first two numbers are the width, the rest denote the length. So 2280 is 22mm x 80mm, and 22110 is 22mm x 110mm.Is there a standard for PSU's about how long they provide power, that SSDs could design to?
Yes. The ATX power supply spec has timing requirements for the PWR_OK signal. As described on Wikipedia:
> The ATX specification requires that the power-good signal ("PWR_OK") [...] remain high for 16 ms after loss of AC power, and fall (to less than 0.4 V) at least 1 ms before the power rails fall out of specification (to 95% of their nominal value).
So power supplies are expected to continue operating through roughly a single missing cycle of AC power, but they only have to give the system 1ms of warning when power is going out. This signal would have to be delivered to the SSD by software in the form of a shutdown notification (a write to a particular NVMe register).
And now that I think about it, that shutdown notification mechanism probably warrants inclusion in the article.
I was hoping for a post-1995, more usable spec since we know that current PSUs often do provide power longer than 1 ms.
A single AC cycle would be 16-20 milliseconds which could already be doable if the SSD got the signal right away without OS involvement.
This potential failure mode is possible when storing more than one bit per flash memory cell, and using a multi-step process to program the cell voltage, and mapping the low-order bits of a cell's value to different LBAs than the high order bit. Drives need to have the capability to either complete or safely abort an in-progress cell program process so that the value in the high-order bit(s) isn't corrupted by an incomplete programming of the low-order bit(s). And this power failure problem isn't the only reason why SSDs need to be careful about leaving cells in a partially-programmed state.
https://www.kingston.com/en/ssd/dc1000b-data-center-boot-ssd
I'm planning to replace it by the SSD linked above.
I was already aware of the power loss issue on consumer SSDs (from the 2013 piece http://lkcl.net/reports/ssd_analysis.html), but I wanted to see it happen for real. I guess I did now.
Could that have been caused by a sudden power loss that corrupted a key section or the drive?
NVMe does not require such a guarantee, nor does it provide a way for drives to signal such a guarantee.
(Part of the reason is that NVMe devices have multiple queues, and the standard tries to avoid imposing unnecessary timing or synchronization requirements between commands that aren't submitted to the same queue.)
1. Issue A and B.
2. Wait for A and B to complete.
3. Issue a FLUSH operation (to ensure that A and B are written from the drive cache to persistent storage), and wait for it to complete.
4. Issue C with FUA (force unit access) bit set.
5. Wait for C to complete.
Alternatively, if the device doesn't support FUA, for writing C you must instead do
4b. Issue C.
5b. Wait for C to complete.
6b. Issue FLUSH, and wait for the FLUSH to complete.
Now, like wtallis already said, NVME additionally has multiple queues per device, but these are independent from each other. If you somehow want ordering between different queues, you must implement that in higher level software.
[1] The SCSI spec has an optional feature to enable ordered tags. But apparently almost no devices ever implemented it, and AFAIK Linux and Windows never use that feature either.
Anecdotally, consumer NVMe SSDs actually tend to not lie about it. Every time I've benchmarked a consumer NVMe SSD under Windows both with and without the "Write Cache Buffer Flushing" option, it has a profound impact on the measured performance of the SSD. I have not observed a comparable performance impact for SATA SSDs, so I suspect Microsoft's description of what that option does is inaccurate for at least one type of drive, though it is at least possible that ignoring flushes is extremely common for consumer SATA SSDs but uncommon for consumer NVMe SSDs.
There's a guy on btrfs' LKML (also the author of [0]) who is diligent enough to do these tests on much of the hardware he gets, and his experience does not sound good for consumer drives.
This isn't quite right. You have to ensure that the drive returned completion of a flush command to the OS before the plug was pulled, or else the NVMe spec does allow the drive to return old data after power is restored. Without confirming receipt of a completion queue entry for a flush command (or equivalent), this test as described is mainly checking whether the drive has a volatile write cache—and there are much easier ways to check that.
TLDR: Very few drives don't implement flush correctly. Notice that he mainly uses hard disks, not SSDs/NVMe. Failure often occurs when two (usually rare) things occur at once. E.g. remapping an unreadable sector while power-cycling.
Went back and found the link. To my surprise, it was 15 years ago. To my greater surprise, the original post, the Slashdot article, and the utility all remain available.
And hard drives (or their NVMe successors) still lie.
Journaling data cuts write throughput in half and it’s not necessary most of the time.
Journaling filesystems can still implement atomic appends with only metadata journaling.
Updating in place is generally not atomic because of the way writeback works for buffered IO.
If you use unbuffered IO you bypass the OS scheduler but still have the disk reordering things if you don’t use write barriers and they still don’t guarantee atomicity for regular writes.
I've found Samsung to be quite a good brand, as well as HP surprisingly. YMMV with HP though.