I tested four NVMe SSDs from four vendors – half lose FLUSH'd data on power loss (2022)
twitter.com
twitter.com
I think the last time I had corrupted files after a power loss was in a FAT32 disk on Win98, but you'd usually get garbage data, not all zeros.
FWIW I've found NTFS and ext3/4 to be of similar reliability over the years, in general use and in the face of improper shutdown. Metadata journaling does a lot to preserve the filesystem in such circumstances. Most of the few significant problems I've had have been due to hardware issues, which few filesystems on their own will help you with.
It is worth noting that when you run tools like chkdsk or fsck, some of the issues reported and fixed are not data damaging, or structurally dangerous, or at least not immediately so. For instance free areas marked in such a way that makes them look used to the allocation algorithms.
However, I'm also not making a statement about NTFS's reliability vs ext3/ext4. In the years that I worked in that position, I maybe dealt with Linux systems 3 times.
Reminder of the original statement:
> NTFS is an extremely resilient yet boring filesystem, I cannot remember the last time I had to run chkdsk even after an improper shutdown.
I never said anything about catastrophic failure of an NTFS filesystem. I have experienced that, but it was comparatively rare (still happened though). I have, however, have had to run chkdsk fairly often to correct errors. Sometimes user data was affected to some degree, but it was often a matter of system stability and getting Windows back to running without issues.
I still find that NTFS is reasonably resilient and have no qualms with it. I just want to push back against the idea that nothing ever goes wrong with NTFS, which was implied by your statement that you can't remember the last time you used chkdsk.
I should have been a bit clearer there: I only mentioned those specific filesystems because those are the ones I have a lot of experience with, rather than intending to bang the drum for them (or one in favour of the other). I expect other much-tested journaled filesystems could be substituted into the same sentences.
> your original statement
GPP was not me.
You are less likely to get garbage with an SSD in combination with a modern filesystem because of TRIM. Even if the SSD has not (yet) wiped the data, it knows that a block that is marked as unused can be retuned as a block of 0s without needing to check what is currently stored for that block.
Traditional drives had no such facility to have blocks marked as unused from their PoV, so they always performed the read and returned what they found which was most likely junk (old data from deleted files that would make sense in another context) though could also be a block of zeros (because that block hadn't been used since the drive had a full format or someone zeroed free-space).
Depending on mount options, ext4fs does metadata journaling ensuring the FS itself is not borked, but not data journaling which would safeguard the file contents in event of unclean shutdown with pending writes in the caches.
The same phenomenon is at play when people complain that their log files contain NUL bytes after a crash. The file system metadata has been updated for the size of the file to fit the appended write, but the data itself was not written out yet.
Getting back zeroes after a metadata sync (which must follow a data sync) would accordingly be an indication of something weird having happened at the disk level: We'd expect to either see no data at all, or correct data, but not zeroes or any other file's or previously written stale data.
I seem to vaguely recall an issue like that, for ext4 in particular. Of course it's possible in general for any filesystem that supports holes, but I don't think we can necessarily assume that the data is always written, and all the pointers to it also written, before the file-size gets updated.
Both extents and the file size are metadata as far as I understand, which would be atomically updated through the journal.
Data can be written before metadata (in data=ordered mode):
> All data are forced directly out to the main file system prior to its metadata being committed to the journal.
Or if it were a hardware ordering fault, remember that SSD TRIM is typically used by modern filesystems to reclaim unused space. TRIMmed blocks read as zero.
Hm, is that a common approach? I thought applications mostly use fallocate(2) for that if it's for performance reasons, which does not change the nominal file size.
Actually allocating zeroes sounds like it could be quite inefficient and confusing, but then again, fallocate is not portable POSIX.
> Or if it were a hardware ordering fault
That's what I suspect might be going on here.
There was a point where ext3 defaulted to data=writeback, which can definitely give you files full of null bytes.
And data=journal exists but is overkill for this situation.
because the only guarantee which data=ordered provides is the security guarantee that stale data won't be revealed.
Yes, it's bad and breaks prefix append consistency, and does not match the documentation...
[0] https://issuetracker.google.com/issues/172227346 [1] https://issuetracker.google.com/issues/172227346#comment8
Also fsync is a terrible API that should be replaced, but that's mostly a different topic.
Or, one can take the ZFS approach and assume the hardware often lies :)
An error is preferred to silently returning garbage data!
Sure, due to their COW nature zfs and btrfs provide better behavior despite broken applications. But you can't solve persistence in the face of lying firmware.
Even thought zfs has some enhancements to not corrupt itself on such drives, if you run for example a database on top, all guarantees around commit go out the window.
Fun fact: rename() is atomic with respect to running applications per POSIX, that the on-disk rename is also atomic is only incidental.
With this rule, three outcomes are acceptable: both occur, or neither occur, or just the file write happens. The unacceptable outcome is that just the rename happens.
("file write" here could mean a single write, or an open-write-close sequence, it doesn't particularly matter and I don't want to dig through old discussions in too much detail)
You mean on complete system crash, right? Your application crashing shouldn't lead to files being fulls of zeroes as long as you've already written everything out.
2/12 is not nearly as dramatic as “half”, and the ones that lost data are the cheap brands as one would expect.
Either choice will lead somebody to complain
Specifically: use the original title, then express your view as a top level comment. If people agree with it, the comment is naturally voted up.
Admittedly this is a twitter thread, so no "actual title" exists.
I'd suggest posting an alternative, non-misleading source instead, or if none exists, don't post at all.
https://hackernewstitles.netlify.app/ has a recent example "Kevin Scott (MSFT CTO) inviting all Open AI employees to join MSFT" which sounds like the title of a third-party article describing said event, but actually directly points to the Microsoft CTO's tweet with the offer https://twitter.com/kevin_scott/status/1726971608706031670 so it got edited to a quote "If needed, you have a role at Microsoft that matches your compensation" which is an attempt to summarize the content but doesn't provide additional context so it might be condidered more clickbaity https://news.ycombinator.com/item?id=38364315
Complain about it in a comment so that people who don't click the link are aware it's misleading. It's not attacking OP.
Is it? I passed on an offer for a drive carrying that name, and got something else for slightly more, the other day as I didn't know the name.
Perhaps their noteworthiness varies internationally? Or do they mainly sell to manufacturers rather than direct to the likes of me?
> Or do they mainly sell to manufacturers rather than direct to the likes of me?
This. Think they mostly sell OEM SSDs under their name, if you buy a laptop or pre built system from a major manufacturer chances are not that low that you find a SK Hynix SSD in there.
(see later in the thread: https://twitter.com/xenadu02/status/1496770658750980103)
The other problem is that this doesn't only cause problems resulting from power loss. At least some files systems guarantee consistency of data on the drive by flushing at critical points. ZFS on Linux does this. If the flush doesn't happen as promised, subsequent writes could result in corrupted files should something else like a crash interrupt operation.
What a sad world to live in, when one comes to expect cheap storage devices to not fulfill intended function.
There were even stories of those crazy coupon people reselling on Amazon, and some cases of returned retail products ending up as “new” on Amazon. Which gets problematic with certain things like consumables (the WSJ did an article on toothpastes iirc).
AMEN to that !
And the most annoying thing is that those of us who know to avoid FBA is that Amazon have removed the "sold by Amazon" search filter tick-box.
So whilst in the past you could tick a box and be presented with a list of products which are direct-sold rather than FBA, you cannot do that anymore.
According to some Reddit posts, you can still do it if you hack the URL and add an "emi=$obscure_value" GET-param. But I'm guessing sooner or later Amazon will kill this work-around too.
Does this actually make a difference? I remember the issue was that Amazon would bin devices together regardless if they're from some random third-party or direct sale, so you could have fakes mixed in with genuine and it was basically a lucky dip.
Is this not still the case?
Ultimately, I'd be weary of buying things like this from Amazon and as you suggest go to a more traditional retail channel instead.
Running dmidecode showed that the part number didn’t match the sticker on the module.
I purchased a bunch of fake kingston SD cards in China that worked well enough for the price, but crapped out within a year of mild use. I didn’t lose data. It was as if one day they worked. Then one day they were fried.
The Samsung Magician app on Windows reports it as "genuine" and it was able to apply two firmware updates. The only thing it complains about is that I should be using PCIE 4 instead of 3, but I can't do anything about that.
The leading theory I have read is that maintenance/refreshing on the ssd is not done preventative/correctly by the firmware and you need to trigger it by accessing the data.
There's some cost cutting somewhere. The NVMEs can't seem to sustain throughput.
It's been pretty disappointing to move I/O bound workloads over and not see notable improvements. The magnitude of data I'm talking about is 500-~3000GB
I've only got two NVME machines for what I'm doing so I'll gladly accept that it's coincidentally flaky bus hardware on two machines, but I haven't been impressed except for the first few seconds.
I know Everyone says otherwise which is why I brought it up. Someone tell me why I'm crazy
Edit: no, I'm not crazy. https://htwingnut.com/2022/03/06/review-leven-2tb-2-5-sata-s... this is similar to what I'm seeing with Crucial and Adata hardware, almost binary performance
Modern platter is actually pretty decent and cheap. It's probably still the way to go for large loads unless you have a grove of money trees
2) Sustaining throughput seems the least of our problems when some unknown number of NVMe SSDs might be literally losing flushed data.
Depending on your needs this can be an issue.
For me, using this drive for random "office" work, I figured I'd never feel it in practice. This drive is supposed to support PCIE 4, but my laptop only does 3. This also "helps", since it won't fill whatever cache it uses as quickly. In practice, it was able to write 100 GB at the top speed. Didn't bother to test more. The only time I've ever written that many data at a time was restoring a backup when I bought the drive. Since my backup was on 2.5" spinning rust, it wasn't an issue.
I'd probably advocate the same.
But some drives are truly awful. My work laptop came with a cheap Samsung drive that would quickly drop to around 200 MB/s. At first, I thought I had somehow badly configured Linux or something (I'm running zfs with native encryption).
Then I went and checked my desktop running quite worn-out SATA ssds from ~2012 (840 evo) and those drives would wipe the floor with the NVMe in write performance. They wouldn't go below 400 something MB/s until almost full. Same kernel version and zfs config.
It would seem that this is quite common behavior in cheap drives. But I guess that if all you do is browse the web, write mails in outlook and type the occasional word document, you're still ahead of spinning rust for the durability (it won't break if you drop your laptop) and for the latency.
Maybe the key is a bunch of small ones. Like 20 512GB modules... That may be brilliant. It's way cheap
However, what I've found perusing those reviews is that there's a huge price gap between this class of drives (9x0 pro, wd 8xx and the like) and "enterprise" drives which seem to have more stable performance.
https://www.tomshardware.com/reviews/samsung-980-m2-nvme-ssd... (https://cdn.mos.cms.futurecdn.net/mzdXcBUJUuxkbYfmQigqQE-970...)
In a 900 second timeframe, the 980 Pro drops down to a bit above 1GB/s, while the worst is somewhere in the tens of MB/s. So there are significant differences in SSD performance profiles which you can only find out about from reviews like this.
Similar results here:
https://www.anandtech.com/show/16504/the-samsung-ssd-980-500...
I can't say I've ever seen a recent SSD (that isn't otherwise faulty) get slow enough to say it is outperformed by a traditional drive, even just counting the fastest end of the disk, but I've certainly seen them drop to around the same speed during a bulk write.
--
[1] unlike this sort of thing: https://www.tomshardware.com/news/low-performance-external-m...
[2] get SLC-only⁴ drives, not QLC-with-SLC-cache or just-QLC, and so forth
[3] bulk data processing tasks such as video editing are where you'll feel this significantly, unless your number-crunching is also bottlenecked at the CPU/GPU
[4] SLC-only is going to be very expensive for large drives, even high-grade enterprise drives tend to be MLC-with SLC-cache. SLC>MLC>TLC>QLC…
[5] this can be quite difficult in the “consumer” market because you'll sometimes find a later revision of the same drive having a completely different memory and/or controller arrangement despite the headline model name/number not changing at all – this is one reason why early reviews can be very misleading
I found some bash script that looped and wrote small blocks synchronously and the HP EX920 was like 20 syncs/sec and my WD RE4 spinner was around 150. Other SSDs were much faster (it was a few years ago so can't remember the exact numbers)
Saw this happen with previous job. I upgraded several Windows devices too Windows 10 and the fastest PC was a Dell desktop with a HDD.
Midrange to lower-mid laptop coupled with low-end SSD's.
(BTW this points out the crappy use of the word “performance” in computing to mean nothing but “speed”. The machine should “perform” what the user requests — if you hired someone to do a task and they didn’t do it, we’d say they failed to perform. That’s what’s going on here.)
The bigger problem is manufacturers chasing the performance. Generally you get the feeling they just hit their firmware with a hammer so it barely doesn't break NTFS.
See also the drama around btrfs' "unreliability", which is all traced back to drives with broken firmware. I fully expect bcachefs will get exactly the same problems.
That's very optimistic of you.
You already have the device shipped to you under $40 with packaging, shipping from China, to the distributor, to you; with a profit at the every step.
I'd gladly pay a few bucks extra for a drive that doesn't delete my data when the power fails
For a $20 device this is 5-25% of the cost the product, assuming $40 retail cost. It would be 1-5% for a $100 product, of course.
You would pay for a such device, but 99% of people wouldn't. Now you have a product which costs 5-25% higher than all your competitors and in the world where the price dictates the sales you have no opportunity to sell your product.
NB to be sure on the prices I've checked Amazon. There are SSDs what are cheaper than many mounting trays, enclosures (without drives) and tray kits.
Whether consumers would actually pay more for a device with power protection is unclear.
So? No reason to break the contract that flush makes all submitted writes durable. The drive can compact space in the background.
I’m pretty sure they used to be on consumer drives too. Then they got removed and all the review sites gave the manufacturer a free pass even though they’re selling products that are inadequate.
Disks have one job, save data. If they can’t do that reliably they’re defective IMO.
Besides, if you're trying to reach the people who remain on Twitter, that says a decent amount about you.
You seem to be claiming that everyone still on Twitter is ideologically compromised. There's a ton of people just ignoring the politics and I still need a channel to reach them.
For example, Twitter is now called X. Surely they noticed that.
Maybe you should stay on Twitter.
Edit: Based on your website, who exactly are you trying to reach? Prospective clients? On Twitter?
You say "users of your app", but it seems like that isn't a huge volume of people, and were you to move to Bluesky, they would almost certainly follow... Unless it's not worth it for them?
So what?
Twitter now being called X can go undetected by users, simply because everything still works the same as before the non-completed namechange.
THAT is what shouldn't be "normalized".
But since you brought it up, of all the indicators of societal descent over the last 7-10 years, the shift away from critical analysis and open dialogue to soviet-level information control and behavior/thought policing is most troubling.
I use Twitter daily because it's the only platform where I have a chance to hear the whole truth on a given subject, usually assembled from multiple sources. I follow a lot of actual leftists as well as the conservatives and liberals-in-exile that make up the modern "right".
I don't know that I follow anyone with whom I agree on every topic, and that's as it should be IMO. Even when I strongly disagree, I'm grateful for the freedom to hear and evaluate these voices.
And while I'm reluctant to engage with your "argument", I'll say that I've seen a handful of antisemitic posts over 10+ years on Twitter, none recently. I have seen a fair amount of anti-Israel and anti-Zionist content in the last six weeks, consistently from those on the left and consistent with what I've seen in demonstrations around the world.
You don’t get to support a hateful platform consequence free.
iirc, twitter turned off all anonymous access, unless you come from a search engine, then you get a limit number of requests. So zedeus came up with an idea to make a massive pool of search engine'd api tokens and use those to keep nitter up. The mirrors would have to copy that idea, and few (0?) have atm.
I think the token workaround you mentioned is the old way that no longer works, but I am not sure.
Correction: “Plus” not “Pro”. Exact model and date codes:
Samsung 970 Evo Plus: MZ-V7S2T0, 2021.10 WD Red: WDS100T1R0C-68BDK0, 04Sept2021
Update 2: models that lost writes: SK Hynix Gold P31 2TB SHGP31-2000GM-2, FW 31060C20 Sabrent Rocket 512 (Phison PH-SBT-RKT-303 controller, no version or date codes listed)
(I always keep an eye on things like firmware updates).
I'm thinking of this example, but also more generally USB devices, Bluetooth devices, etc.
https://ftdichip.com/wp-content/uploads/2020/08/TN_114_USB-D...
https://www.usb.org/sites/default/files/usb-if_original_logo...
It might be a different story if the spec was intentionally violated, though (rather than incidentally, i.e. due to an idea that should have been transparent/indistinguishable externally but didn't work out).
It's their responsibility to do develop the product correctly, do QA, and if a defect is found, advise customers or stop selling the defective goods.
The greatest scam the computer industry pulled was convincing people that computers are magical, unpredictable devices that are too complex for the industry to be held responsible for things not working as claimed.
I do agree that there should be transparency, though: Label it “turbo mode”, add a big warning sticker (and ideally a way to opt out of it via software or hardware), but don’t just pretend to be able to have the cake and eat it too.
For extra fun: if the box carries a trademark from a standards group, you could try adding them into the suit; use of their trademarked logo could be argued to be implied fitness, if there are standards the drive is supposed to meet to use it.
At the very least they might get tired of the expense of the expense of sending someone to defend the claim, and it would cease to be profitable to engage in this scammery.
Always makes me laugh.
Anyways, not in the US where you're probably asking for, but yes, the vast majority of the developed word has that. It's called "false advertising", and exists at least in the EU, Australia, UK. You can't put a label on your product or advert that is false or misleading.
So if the box says this is a WiFi6E router, but it's actually only 5 because it's using the wrong components to save on costs, you can report them to the relevant authority and they'll be fined (and depending on the case and scenario you get compensation). The process is a harder bordering on the impossible if you bought from AliExpress from a random no name vendor though, but as long as the vendor or platform or store exists in the country with the sensible regulation you can report it.
I think the question is less “if they skip on parts and lie” and more along the lines of incompleteness. Like “it’s an HTTP server, but they saved on effort and implement put as post, which works fine for most of use cases”.
That said, I’d guess this would be a pretty hard case to win. The law typically requires intent when false advertising, so if they didn’t know they didn’t follow the spec they might be fine. And it depends on the claims and what the consumer can expect. Like, if you deliberately don’t say explain the exact spec your SSD complies with, and you make no explicit promises of compatibility, it’s a harder win. Like I bet few SSD manufacturers will say “Serial ATA v3.5 (may 2023) tested and compatible with OpenXFS commit XYZ on Debian Linux running kernel version 4.3.2”. But if they say “super fast SSD with a physical SATA cable socket”, then what really was false if it doesn’t support the full spec?
Wondering if anything changed since the original tests...
You're wondering if firmware writers lie to layers higher up in the stack? I think it's a 100% certainly that there's drive firmware that lies.
There's a reason why many vendors have compatibility lists, approved firmware versions, and even their "own" (rebranded from an OEM) drives that you have to buy if you want official support (and it's not entirely a money grab: a QA testing infrastructure does cost money).
I have very little trust in consumer flash these days after seeing the firmware shortcuts and stealth hardware replacements manufacturers resort to to cut costs.
I wonder how much perf is on the table in various scenarios when we can give up needing to flush. If you know the drive has some resilience, say, 0.5s of time it can safely writeback during, maybe you can give up flushes (in some cases). How much faster is the app then?
It's be neat to see some low-cost improvements here. Obviously in most cases, just get an enterprise drive with supercapa or batteries onboard. But an ATX power rail that has extra resilience from the supply, or an add-in/pass-through 6-pin sata power supercap... that could be useful too.
Flush+FUA requires the data to be stored to non-volatile media. Capacitor-backed RAM dumping to flash is non-volatile. When a drive knows it has enough capacitor-time to finish flushing all preceding writes from the cache, it can immediately say the flush was completed. This can all be handled on the device without the software having to make guesses at how long something has to be written before it's durable.
During typical usage the flash controller is constantly journaling LBA to physical addresses in the background, so that the entire logical to physical table isn’t lost when the drive loses power. With a larger capacitor you could potentially remove this background process and instead flush the entire logical to physical table when the drive registers power loss. But as this area makes up ~2% of the total NAND, that’s at absolute best a 2% performance benefit we are potentially missing out on.
This is not a technical problem that needs yet another SATA/SAS/etc command to be standardized. It's a 'social' problem that there's no real incentives for firmware writers to tell the truth 100% of the time.
The best you can hope for is if you buy a fancy-pants enterprise storage solution with compatibility lists and approved firmware versions.
Not even OS vendors are immune to the temptations of higher performance through somewhat relaxed interpretation of interfaces: https://developer.apple.com/library/archive/documentation/Sy...
After that command it should be 100% safe to pull the power because everything SHOULD have been written to flash. That’s the point of the command.
It’s interesting that the drives that do it wrong still take time indicating they’re doing something.
Especially for consumer hardware on Linux--there's a lot of stuff that "works" but is not necessarily stable long term or that required a lot of hacking on the kernel side to work around issues
Did they Pass or Fail the flush test? (Saved or Lost data respectively?)
> SK Hynix Gold P31 2TB SHGP31-2000GM-2, FW 31060C20
> Sabrent Rocket 512 (Phison PH-SBT-RKT-303 controller, no version or date codes listed)
Looks like those two models failed.
Would like to note that these tweets are from Feb 22, 2022. This entire thread should have (2022) on it.
Samsung 970 Evo Plus: MZ-V7S2T0, 2021.10: Pass
WD Red: WDS100T1R0C-68BDK0, 04Sept2021: Pass
Crucial P2 250GB CT250P2SSD8, FW P2CR046: Pass
Samsung 980 250GB MZ-V8V250, 2021/11/07: Pass
WD Black SN750 1TB WDS100T1B0E, 09Jan2022: Pass
WD Green SN350 240GB WDS240G20C, 02Aug2021: Pass
SK Hynix Gold P31 2TB SHGP31-2000GM-2, FW 31060C20: Fail
Sabrent Rocket 512 (Phison PH-SBT-RKT-303 controller, no version or date codes listed): Fail
I agree about Sabrent.
On a related note I tested 4 DDR5 Ram kits from major vendors - half of them corrupt data when exposed to UV light.
I just need to know when the data is written and when it isn't.
[1] https://www.truenas.com/community/threads/x4-pcie-to-nvme-ad...
"Buy cheap, buy twice" as they say... =)
The bottom answer especially states that "blockdev --flushbufs may still be required if there is a large write cache and you're disconnecting the device immediately after"
The hpdarm utility has a parameter for syncing and flushing device buffers themselves. Seems like all three should be done for a complete flush at all levels.
The rule is not that hard to remember.
I am sure it might be easy to see visually - a lack of substantial capacitor on the board would indicate a high likelihood of data loss.
Not scenarios where the data is in the cache when the power drops and the drive is expected to write out the cache with internal power alone.
The first scenario is for consumer disks, they don't report 'committed' until the data is actually committed. That makes them a whole assload slower. (Unless they lie.)
Looking up the definitions of write-back vs write-thru caching will help you out here
6.8 Flush command
The Flush command shall commit data and metadata associated with the specified namespace(s) to non-volatile media. The flush applies to all commands completed prior to the submission of the Flush command.