I tested four NVMe SSDs from four vendors – half lose FLUSH’d data on power loss
twitter.com
twitter.com
It would be nice if we could just buy dumb flash and let the application do whatever it wants (I guess that application would be your filesystem; but it could also be direct access for specialized use cases like databases). If you want maximum speed, adjust your settings for that. If you want maximum write durability, adjust your settings for that. People are always looking for that one size fits all use case, but it's hard here. Some people may be running cloud providers and already have software to store that block on 3 different continents. Some people may be an embedded system with a fixed disk image that changes once a year, with some temporary storage for logs. There probably isn't a single setting that gets optimal use out of the flash memory for both use cases. The cloud provider doesn't care if a block, flash chip, drive, server rack, availability zone, or continent goes away. The embedded system may be happy to lose logs in exchange for having enough writes left to install the next security update.
It's all a mess, but the constraints have changed since we made the mess. You used to be happy to get 1/6th of a PCI Express lane for all your storage. Now processors directly expose 128 PCIe lanes and have a multitude of underused efficiency cores waiting to be used. Maybe we could do all the "smart" stuff in the OS and application code, and just attach commodity dumb flash chips to our computer.
1. Contemporary mainstram OSes have not risen to the challenge of dealing appropriately with the multi-CPU, multi-address space nature of modern computers. The proportion of the computer that the "OS" runs on has been shrinking for a long time and there have only been a few efforts to try to fix that (e.g. HarmonyOS, nrk, RTKit)
2. Hardware vendors, faced with proprietary or non-malleable OSes and incentives to keep as much magic in the firmware as possible, have moved forward by essentially sandboxing the user OS behind a compatibility shim. And because it works well enough, OS developers do not feel the need to adjust to the hardware, continuing the cycle.
There is one notable recent exception in adjusting filesystems to SMR/Zoned devices. However this is only on Linux, so desktop PC component vendors do not care. (Quite the opposite: they disable the feature on desktop hardware for market segmentation)
When this structure is fed to a EDR/HDR/FDR Infiniband network, the result is a blazing fast storage system where you can make a massive number of random accesses by very large number of servers simultaneously. The whole structure won't shiver even.
There are also other tricks Lustre can pull for smaller files to accelerate the access and reduce the overhead even further, too.
In this model, the storage boxes are somewhat sandboxed, but the whole model as a general is mounted via its own client, so the OS is very close to the model Lustre provides.
On the GPU servers, if you're going to provide big NVMe scratch spaces (a-la nVidia DGX systems), you soft-RAID the internal NVMe disks with mdadm.
In both models, saturation happens on hardware level (disks, network, etc.) processors and other soft components doesn't impose a meaningful bottleneck even under high load.
Also, yes, many longer jobs are checkpoints and restart where it's left off, but it's not always possible, unfortunately.
Or IBM's GPFS / Spectrum Scale. Same deal, really, although GPFS is a more complete package.
It explores how modern systems are a set of cooperating devices (each with their own OS) while our main operating systems still pretend to be fully in charge.
Hence why SSD's use a block layer (or in the case of NVMe key/value, hello 1964/CKD) abstraction above whatever pile of physical flash, caches, non-volatile caps/batts, etc. That abstraction holds from the lowliest SD card, to huge NVMe-OF/FC/etc smart arrays which are thin provisioning, deduplicating, replicating, snapshoting, etc. You wouldn't want this running on the main cores for performance and power efficiency reasons. Modern m.2/SATA SSD's have a handful of CPUs managing all this complexity, along with background scrubbing, error correction, etc so you would be talking fully heterogeneous OS kernels knowledgeable about multiple architectures, etc.
Basically it would be insanity.
SSDs took a couple orders of magnitude off the time values of spinning rust/arrays, but many of the optimization points of spinning rust still apply. Pack your buffers and submit large contiguous read/write accesses, queue a lot of commands in parallel, etc.
So, the fundamental abstraction still holds true.
And this is true for most other parts of the computer as well. Just talking to a keyboard involves multiple microcontrollers, scheduling the USB bus, packet switching, and serializing/deserializing the USB packets, etc. This is also why every modern CPU has a mgmt CPU that bootstraps and manages it power/voltage/clock/thermals.
So, hardware abstractions are just as useful as software abstractions like huge process address spaces, file IO, etc.
I honestly don’t know where you got that idea from. I always thought the whole point of computer programming was to solve problems. If it makes things more complex as a result, then so be it. Just as long as it creates fewer, less severe problems than it solves.
And this is where the increased complexity is necessary for a solution, not just Perl Anti-golf or FactoryFactory Java jokes.
I don't think you can get away from this - yes, can solve a problem, but if you model problems as entropy, increasing complexity increases entropy.
It's like the messy room problem - you can clean your room (arguably high entropy state), but unless you are exceedingly careful doing so increases entropy. You merely move whatever mess to the garbage bin, expend extra heat, increase your consumption in your diet, possibly break your ankle, stress your muscles...
Arguably cleaning your room is important, but decreasing entropy? That's not a problem that's solvable, not in this universe.
In an isolated system, entropy can only increase. Moving at all heats up the air. Even if you are exceedingly careful, you increase entropy when doing any useful work.
But, its still an abstraction, and would have to remain that way unless your willing to segment it into a bunch of individual product categories, since the functionality of these controllers grows with the target market. AKA the controller on a emmc isn't anywhere similar to the controller on a flash array. So like GP-GPU programming, its not going to be a single piece of code because its going to have to be tuned to each controller, for perf as well as power reasons never mind functionality differences (aka it would be hard to do IP/network based replication if the target doesn't have a network interface).
There isn't anything particularly wrong with the current HW abstraction points.
This "cheating" by failing to implement the spec as expected isn't a problem that will be solved by moving the abstraction somewhere else, someone will just buffer write page and then fail to provide non volatile ram after claiming its non volatile/whatever.
And its entirely possible to build "open" disk controllers, but its not as sexy as creating a new processor architectures. Meaning RISC-V has the same problems, if you want to talk to industry standard devices (aka the NVMe drive you plug into the RISC-V machine is still running closed source firmware, on a bunch of undocumented hardware).
But take: https://opencores.org/projects?expanded=Communication%20cont... for example...
That’s why I suggested something similar to OpenFirmware. With that in place, you could send a piece of Forth code and the controller would run it without involving the CPU or touching any bus other than the internal bus in the storage device.
Plus, disks are at least partially sold on their performance profile, and your going create another problem with a cross platform IR vs coding it in C/etc and using a heavyweight optimizing compiler targeting the microcontroller in question directly.
You also have to remember these are effectively deeply embedded systems in many cases, which are expected to be available before the OS even starts. Sometimes that includes operating over a network of some form (NVMe-OF). So it doesn't even make sense when that network device is shared/distributed.
I'd probably pay for "unlocking" ZNS on my Samsung 980 Pro, if just to reduce the write amplification.
From what I understand the abstraction works a lot like virtual memory. The drive shows up as a virtual address space pretending to be a a disk drive and then the drive's firmware maps virtual addresses to physical ones.
That doesn't seem at all incompatible with exposing the mappings to the OS through newer APIs so the OS can inspect or change the mappings instead of having the firmware do it.
The problem is that it's not a good match for how flash memory actually works, especially with regards to the extreme disparity between a NAND page write and a NAND erase block. Giving the OS an interface to query which blocks the SSD considers as live/allocated rather than deallocated and implicitly zero doesn't seem all that useful. Giving the OS an interface to manipulate the SSD's logical to physical mappings (while retaining the rest of the abstraction's features) would be rather impractical, as both the SSD and the host system would have to care about implementation details like wear leveling.
Going beyond the current HDD-like abstraction augmented with optional hints to an abstraction that is actually more efficient and a better match for the fundamental characteristics of NAND flash memory requires moving away from a RAM/VM-like model and toward something that imposes extra constraints that the host software must obey (eg. append-only zones). Those constraints are what breaks compatibility with existing software.
There's a _lot_ of problems doing high performance NAND yourself. You honestly don't want to do that in your app. If vendors would provide full specs and characterization of NAND and create software-suitable interfaces for the device then maybe it would be feasible to do in a library or kernel driver, but even then it's pretty thankless work.
You almost certainly want to just buy a reliable device.
Endurance management is very complicated. It's not just a matter of PE cycles for any given block will meet UBER spec at data retention limits with the given ECC scheme. Well, it could be in a naive scheme but then your costs go up.
Even something as simple as error correction is not. Error correction is too slow to do on the host for most IOs, so you need hardware ECC engines on the controller. But those become very large if you have a huge amount of correction capability in them so if errors exceed their capability you might go to firmware. Either way, the error rate is still important to know the health of the data, so you would need error rate data to be sent side-band with the data by the controller somehow. If you get a high error rate, does that mean the block is bad or does it mean you chose the wrong Vt to issue the read with, retention limit was approached, the page had read disturb events, dwell time was suboptimal, operating temperature was too low? All these things might factor in to your garbage collection and endurance management strategy.
Oh and all these things depend on every NAND design/process from each NAND manufacturer.
And then there's higher level redundancy than just per-cell (e.g., word line, chip, block, etc). Which all depend on the exact geometry of the NAND and how the controller wires them up.
I think better would be a higher level logical program/free model that sits above the low level UBER guarantees. GC would have to heed direction coming back from the device about what blocks must be freed, and what the next blocks to be allocated must be.
I don't know, maybe if there was a "my files exist after the power goes out" column on the website, then I'd sort by that, too?
(And they could review memory, too, and do backgrounder videos about standards and commonly available parts.)
more like, "don't lose the last 5 seconds of writes if the power goes out". If ordering is preserved you should keep your filesystem, just lose more writes than you expected.
But then again, if they're willing to accept and confirm flush commands without flushing, I wouldn't expect them to actually follow ordering constraints.
you can use some sort of WAL mechanism to ensure that the the parallel writes appear as if ordering was preserved. that will allow you to lie and ignore fsyncs, but still ensure consistency in case of a crash.
>But then again, if they're willing to accept and confirm flush commands without flushing, I wouldn't expect them to actually follow ordering constraints.
it depends on which type of liar they think you are. if they're the "don't care, disable all safeguards type", then yes they're probably ignoring ordering as well. However, it's also possible they're the methodical liar, figuring out what they can get away with. As I mentioned in another comment in this thread, as long as ordering is preserved the lie wouldn't be noticed in typical use cases (ie. not using it for some sort of prod db, and not using it as part of a multi-drive array). power losses are relatively common, and a drive that totally corrupts the filesystem on it will get noticed much more quickly than a drive that merely loses the last few seconds of writes.
There is sometimes a whole device erase function provided, but it turns out that a significant portion of tested devices don't actually manage to do that.
Computerphile made a pretty good video about TPMs: https://www.youtube.com/watch?v=RW2zHvVO09g
Are there hardware or OS approaches that facilitate this?
Yeah, this is how things are supposed to be done and the fact it's not happening is a huge problem. Hardware makers isolate our operating systems in the exact same way operating systems isolate our processes. The operating system is not really in charge of anything, the hardware just gives it an illusory sanboxed machine to play around in, a machine that doesn't even reflect what hardware truly looks like. The real computers are all hidden and programmed by proprietery firmware.
I haven't done research on flash design in almost a decade back when I worked on backup software, and my conclusions back then were basically that: you're just better off buying a reliable drive that can meet your your own reliability/performance characteristics, and making tweaks to your application to match the underlying drive operational behavior (coalesce writes, append as much as you can, take care with multithreading vs HDDs/SSDs, et cetera), and testing the hell out of that with a blessed software stack. So we also did extensive tests on what host filesystems, kernel versions, etc seemed "valid" or "good". It wasn't easy.
The amount of complexity to manage error correction and wear leveling on these devices alone, including auxiliary constraints, probably rivals the entire Linux I/O stack. And it's all incredibly vendor specific in the extreme. An auxiliary case e.g. the case of the OP, of handling power loss and flushing correctly, is vastly easier when you only consider some controller firmware and some capacitors on the drive, versus the whole OS being guaranteed to handle any given state the drive might be in, with adequate backup power, at time of failure, for any vendor and any device class. You'll inevitably conclude the drive is the better place to do this job precisely because it eliminates a massive amount of variables like this.
"Oh, but what about error correction and all that? Wouldn't that be better handled by the OS?" I don't know. What do you think "error correction" means for a flash drive? Every PHY on your computer for almost every moderately high-speed interface has a built in error correction layer to account for introduced channel noise, in theory no different than "error correction" on SSDs in the large, but nobody here is like, "damn, I need every number on the USB PHY controller on my mobo so that I can handle the error correction myself in the host software", because that would be insane for most of the same reasons and nearly impossible to handle for every class of device. Many "Errors" are transients that are expected in normal operation, actually, aside from the extra fact you couldn't do ECC on the host CPU for most high speed interfaces. Good luck doing ECC across 8x NVMe drives when that has to go over the bus to the CPU to get anything done...
You think you want this job. You do not want this job. And we all believe we could handle this job because all the complexity is hidden well enough and oiled by enough blood, sweat, and tears, to meet most reasonable use cases.
Apple still has a flash storage controller that exists entirely separately from the host CPU, and the host software stack, just like all existing flash drives do today. The difference? The controller just doesn't exist on the literal, physical drive next to the flash chips. Because it doesn't exist; they just solder flash directly on the board without a mount like an M.2 drive. Again, no variability here, so it can all be "hard coded." And the storage controller instead exists by the CPU in the "T2 security chip", which also handles things like in-line encryption on the path from the host to the flash (which is instead normally handled by host software, before being put on the bus). It also does some other stuff.
So it's not a matter of "architecture", really. The architecture of all T2 Macs which feature this design is very close, at a high level, to any existing flash drive. It's just that Apple is able to put the NVMe controller in a different spot, and their "NVMe controller" actually does more than that; it doesn't have to be located on a separate PCB next to the flash chips at all because it's not a general 3rd party drive. It just has to exist "on the path" between the flash chips and the host CPU.
And it breaks down the instant you change the hardware, even in the slightest ways. Frequently the optimizations then made turn around and reduce the speed below naive methods. Modern flash+controllers are massively more complex than the old NOR flash of two decades ago. Which is why they get multiple CPUs managing them.
It does need support from the storage device though.
Clearly the vendors are at odds with the law, selling a storage device that doesn't store.
I think they are selling snake-oil, otherwise known as commiting fraud. Maybe they made a mistake in design, and at the very least they should be forced to recall faulty products. If they know about the problem and this behaviour continues ait is basically a fraud.
We allow this to continue, and the manufacturers that actually do fulfill their obligations to the customer suffer financially, while unscurpulous ones laugh all the way to the bank.
And as a result, we have an entire generation of machines that cannot ever be trusted. And an awful lot of people seem fine with that, or just haven't fully considered what it implies.
Which is currently all of them, but that was also the case in the distributed systems space when he first started working on Jepsen.
And at the point where that documentation had been done, that on its own might be enough to right the ship without anyone actually having to get sued.
The tester is running the device out of spec.
The manufacturers warrant these devices to behave on a motherboard with proper power hold up times, not in whatever enclosures.
If the enclosure vendor suggests that behavior on cable pull will fully mimick motherboard atx power loss then that is fraud. But they probably have fine print about that, I'd hope.
Thats an interesting point, doesn't 'power failure' also include potential failure of the power supply, in which case you might not get that time?
Or what if a new write command is issued withing the holdup time, does the motherboard /OS know about powerloss during those 16 milliseconds that the power is still holding?
Anyway, let's firm up how an SSD works and what the OS knows.
SSDs have volatile DRAM buffers as a staging area to use before writing to the flash.
Flush (OS ioctl) means the data is successfully residing in the volatile DRAM of the SSD.
This is all the OS knows and usually ever knows in the ioctl cycle.
If power is lost there is some time before the >16ms is up that power good signal is lost on the motherboard. The voltage on the 3.3V rail will probably also drop enough from nominal to let the SSD controller know it better gets its housekeeping in order. In other words, dump the DRAM somewhere permanent and deal with it on the next power up.
Anything the OS is doing in the interim will not likely be acknowledged as flushed so that's not a concern. The OS userspace write will never complete. That loop works fine.
The thing that gets people up in arms is that flush means the SSD has the data only in volatile memory and not necessarily in non-volatile storage.
All performant SSDs seem to work this way. They need buffers.
The larger form factor enterprise drives, which are maybe 25% more expensive, have PLP capacitor banks. These supply a solid 50ms of power. Some manufacturers supply oscilliscope screenshots and such.
Anything else seems to be variable in its approach to power loss, particularly the smaller, hotter M.2 parts .
Capacitor banks have issues like taking up space, causing inrush currents, gaining impedance over time, and mediocre reliability at the high temperatures that latest M.2 sticks experience.
I wonder if there would be a market for a small board that contains the capacitors and passes the signals down to a M.2 female connector. The physical disconnect would probably help with the temperature as well (and the board could come with its own heatsink).
> just attach commodity dumb flash chips to our computer
I kind of agree with your stance; it would be nice for kernel- or user-level software to get low-level access to hardware devices to manage them as they see fit, for the reasons you stated.
Sadly, the trend has been going toward smart devices for a very long time now. In the very old days, stuff like floppy disk seeks and sector management were done by the CPU, and "low-level formatting" actually meant something. Decades ago, IDE HDDs became common, LBA addressing became the norm, and the main CPU cannot know about disk geometry anymore.
https://vadosware.io/post/everything-ive-seen-on-optimizing-...
Granted I was doing something highly questionable (running postgres with fsync off on ZFS) It was very painful to get to the actual issue, but I'm glad I found out.
I've always wondered if it was worth pursuing to start a simple data product with tests like these on various cloud providers to know where these corners are and what you're really getting for the money (or lack thereof).
[EDIT] To save people some time (that post is long), the command to set the feature is the following:
nvme set-feature -f 6 -v 0 /dev/nvme0n1
The docs for `nvme` (nvme-cli package, if you're Ubuntu based) can be pieced together across some man pages:https://man.archlinux.org/man/nvme.1
https://man.archlinux.org/man/nvme-set-feature.1.en
It's a bit hard to find all the NVMe features but 6 is the one for controlling write-back caching.
https://unix.stackexchange.com/questions/472211/list-feature...
[1]https://github.com/linux-nvme/nvme-cli/blob/master/nvme-prin...
[2]https://github.com/linux-nvme/libnvme/blob/master/src/nvme/t...
Unlike eg. ATA and SCSI, the NVMe specs are freely available to the public. They're a little more complicated to read now that the spec has been split into a few modules, but finding the descriptions of all the optional features isn't too hard.
I can't say I'm going to read the spec any time soon but thanks for sharing this pointer, I'll refer here.
Still would be nice to have some of this information in the man page though...
If fsync works and fdatasync does not, that strongly suggests a kernel or filesystem bug in the implementation of fdatasync that should be fixed.
That said, I looked at the logs you showed, and those "Bad Address" errors are the EFAULT error, which only occurs in buggy software, or some issue with memory-mapping. I don't think you can conclude that NVMe writes are going missing when the pg software is having EFAULTs, even if turning off the NVMe write cache makes those errors go away. It seems likely that that's just changing the timing of whatever is triggering the EFAULTs in pgbench.
Yeah I thought the same initially which is why I was super confused --
> If fsync works and fdatasync does not, that strongly suggests a kernel or filesystem bug in the implementation of fdatasync that should be fixed.
Gulp.
> That said, I looked at the logs you showed, and those "Bad Address" errors are the EFAULT error, which only occurs in buggy software, or some issue with memory-mapping. I don't think you can conclude that NVMe writes are going missing when the pg software is having EFAULTs, even if turning off the NVMe write cache makes those errors go away. It seems likely that that's just changing the timing of whatever is triggering the EFAULTs in pgbench.
It looks like I'm going to have to do some more experimentation on this -- maybe I'll get a fresh machine and try to reproduce this issue again.
What led me to NVMe as dropping write was the complete lack of errors on the pg and OS side (dmesg, etc).
I loved seeing how giddy Linus got while testing Valve's Steamdeck, but when it comes to actual benchmarks and testing, I would appreciate if they dropped the entertainment aspect.
I feel like they've successfully developed enough clout/trust that they have escaped the hell of having to pull punches in order to assure they get review samples.
They eviscerated AMD for the 6500xt. They called out NZXT repeatedly for a case that was a fire hazard (!). Most recently they've been kicking Newegg in the teeth for trying to scam them over a damaged CPU. They've called out some really overpriced CPU coolers that underperform compared to $25-30 coolers. Etc.
I bet they'd go for testing this sort of thing, if they haven't already started working on it already. They'd test it and then describe for what use cases it would be unlikely to be a problem vs what cases would be fine. For example, a game-file-only drive where if there's an error you can just verify the game files via the store application. Or a laptop that's not overclocked and only is used by someone to surf the web and maybe check their email.
> for starters i think the lab is going to focus on written for its own content and then supporting our other content [mainly their unboxing videos]... or we will create a lab channel that we just don't even worry about any kind of upload frequency optimization and we just have way more basic, less opinionated videos, that are just 'here is everything you need to know about it' in video form if, for whatever reason, you prefer to watch a video compared to reading an article
https://linustechtips.com/topic/1410081-valve-left-me-unsupe...
I think the current industry players have tunnel vision and are too focused on their balance sheets. Things like reputation, trust, and goodwill are crucial to their businesses, but no one is getting a bonus for something that doesn't directly translate into revenue, so those things get ignored. That kind of short sighted thinking has left the whole industry vulnerable to up and coming influencers who have more incentive to care about things like reputation and brand loyalty.
I've been watching LTT with a fair bit of interest to see if they can come up with a winning formula. The biggest problem is that in-depth technical analysis isn't exciting. I remember reading something many years ago, maybe from JonnyGuru, where the person was explaining how most visitors read the intro and conclusion of an article and barely anyone reads the actual review.
Basically you need someone with a long term vision who understands the value you get from in-depth technical analysis and doesn't care if the cost of it looks bad on the balance sheet. Just consider it the cost of revenue for creating content and selling merchandise.
The most interesting thing with LTT is that I think they've got the pieces to make it work. They could put the most relevant / interesting parts of a review on YouTube and skew it towards the entertainment side of things. Videos with in-depth technical analysis could be very formulaic to increase predictability and reduce production costs and could be monetized directly on FloatPlane.
That way they build their own credibility for their shallow / entertaining videos without boring the core audience, but they can get some cost recovery and monetization from people that are willing to pay for in-depth technical analysis.
I also think it could make sense as bait to get bought out. If they start cutting into the traditional review industry someone might come along and offer to buy them as a defensive move. I wonder if Linus could resist the temptation of a large buyout offer. I think that would instantly torpedo their brand, but you never know.
https://www.rtings.com/monitor/tools/table
They rigorously test their hardware and you can filter/sort by literally hundreds of stats.
I just built a PC and I would have killed for a site that had apples-to-apples benchmarks for SSDs/RAM/etc. Motherboard reviews especially are a huge joke. We're badly missing a site like that for PC components.
Userbenchmark has benchmarks for SSDs[1] and RAM[2]. Can't help you with motherboards though.
Worse yet every time I benchmark one of my machines, I score significantly higher than the average user results for the same hardware. Perhaps the average submitter has crapware/antivirus installed or their machines are misconfigured (e.g. XMP disabled) which makes all their data suspect.
fwiw the best motherboard comparison I found was on overclock.net[1]. It didn't list everything I cared about, but it was a great starting point
[1]: https://www.overclock.net/threads/official-intel-z690-mother...
> apparently WD released an NVMe drive that's faster than Intel Optane
"Faster" is a matter of opinion; it depends on what you're optimizing for. Optane obviously has faster random reads, but it's not so great at sequential writes. The UserBenchmark score tries to take all of that into account: https://ssd.userbenchmark.com/Compare/Intel-905P-Optane-NVMe...
Regarding influencers: they're being leveraged by companies precisely because they are about "the experience", not actual subjective analysis and testing. 99% of the "influencers" or "digital content creators" don't even pretend to try to do analysis or testing, and those that do generally zero in on one specific, usually irrelevant, thing to test.
How about Tom's Hardware and AnandTech? If they don't count, who does? Years ago, I used to read CDRLabs for optical media drives. Their reviews were very scientific and consistent. (Of course, optical media is all but dead now.)
Like... what if LTT bought out Anandtech instead of trying to spin up a new 'labs' to replicate what has largely been lost (but still exists to an extent) in a few dusty corners of the tech journalism world.
I'm willing to give the benefit of the doubt, but there's so far been a lot of "just try me" and "we hired someone amazing" but I'll believe it when I see results!
1. https://www.anandtech.com/show/13092/future-plc-to-acquire-c...
2. https://ca.finance.yahoo.com/quote/FUTR.L?p=FUTR.L&.tsrc=fin...
So if LTT does start providing more objective benchmarks and reviews it could be a powerful market force.
Linus is welcome to chase whatever niche market he wants, but as a "purely informational source" he's got a pretty marred track record these days.
I'm also very curious to see a source on the "bribes he received from Nvidia/Intel", because I'm not finding anything that looks relevant on Google.
Take for example this recent video: "We ACTUALLY downloaded more RAM" [1] complete with grinning youtube face holding a stick of RAM marked '10TB' - and it's complete bullshit.
How can I trust the opinion of someone who publishes such embarrassing nonsense?
If LTT can succeed and provide good information, then what is wrong with their click bait thumbnails (or similar techniques)?
Not sure about his other stuff you claim, I'm not a super big video guy for tech things (just let me skip and search ahead easily) but this came up with my friends a few years ago after people started noticing many videos from various creators going to this format of thumbnail.
Edit: Here is an old paper on the subject of OS filesystem error handling.
https://research.cs.wisc.edu/wind/Publications/iron-sosp05.p...
> The models that never lost data: Samsung 970 EVO Pro 2TB and WD Red SN700 1TB.
I always buy the EVO Pro’s for external drives and use TB to NVMe bridges and they are pretty good.
Samsung have removed the Elpis controller from the 980 PRO and replaced it with an unknown one, and also removed any speed reference from the spec sheet.
Take a look here for what's changed on the 980 PRO: https://www.guru3d.com/index.php?ct=news&action=file&id=4489...
It's OK for them to do this, but then they should give the new product a new name, not re-use the old name so that buying it becomes a "silicon lottery" as far as performance goes.
I knew about the changed controller in the 970 Evo Plus, but I wasn't aware they also changed the 980 Pro. That's disappointing. Is there anyone left that isn't doing those shenanigans?
But IIRC Samsung was also called out for switching controllers last year.
"Yes, Samsung Is Swapping SSD Parts Too | Tom's Hardware"
> Correction: “Plus” not “Pro”. Exact model and date codes:
> Samsung 970 Evo Plus: MZ-V7S2T0, 2021.10 > WD Red: WDS100T1R0C-68BDK0, 04Sept2021
Samsung 970 Evo Plus: MZ-V7S2T0, 2021.10
WD Red: WDS100T1R0C-68BDK0, 04Sept2021
The new Micron 7400 Pro M.2 960GB is $200, for example.
Sure, the published IOPS figures are nothing to write home about, but drives like these 1) hit their numbers every time, in every condition, and 2) can just skip flushes altogether, making them much faster in uses where data integrity is important (and flushes would otherwise be issued).
https://news.ycombinator.com/item?id=30371857
The Samsung EVO drives are interesting because they have a few GB of SLC that they use as a secondary buffer before they reflush to the MLC.
I'm nitpicking, but an EVO has TLC. Also an SLC write cache is the norm for any high performance consumer ssd, it's not just Samsung.
b...but the M in MLC stands for multi... as in multiple... right?
checks
Oh... uh; Apparently the obvious catch-all term MLC actually only refers to dual layer cells, but they didn't call it DLC, and now there's no catch-all term for > SLC. TIL.
Interestingly enough, Apple has a patent for 8-bit cell flash (256 levels!), going full blown analog processing and error correction, but I don't think that ever became a product.
Even Samsung's last MLC drives (970 Pro?) had SLC write caches. Otherwise they'd just be giving up a benchmark.
Enterprise is a whole different game tho, as the value of an SLC cache is inversely proportional to the write duty cycle...
This does a disservice to those who might be running drives from those vendors with an expectation that they don't lose data post-flush.
That said, this narrows one of the data losers down to Hynix. Curious about the other one, considering how many US-based SSD vendors there are.
Not really. Samsung builds a plethora of SSDs.
I didn't pay careful attention to the wording of the submitted title. I may have been confused because of the wording of the actual tweet: I tested a random selection of four NVMe SSDs from four vendors.
The word "random" meant to me that Samsung drives could have been selected twice. But, yes, then there wouldn't be four distinct vendors.
Unstated but implied by you is there are only two (major) Korean vendors to choose from.
So if Samsung is a Korean winner then Hynix must be the Korean loser. Which is now clear to me.
Is it possible there's a third (minor) Korean player? Could I possibly still have a chance? :)
Read documents and specifications like this tester didn't do.
And don't use random enclosures and pull the plug since the design spec assumes hold up times and sequencing that enclosure may not be compliant with.
The NVMe spec is available for free; you should read it.
And you're 100% wrong about the enclosure too. It's driven by an Intel TB bridge JHL6240 and the drives are PCIe NVMe m.2 devices. Power specs are identical to on-board m.2 slots with PCIe support (which is all modern ones). There is no USB involved.
See my other reply to you where I explain what Flush actually does (your comments about it are also completely wrong).
Your TB test sounds valid but did you verify with manufacturer that power loss protection or power failure protection works in your TB enclosure? Is that a fair assumption or do you need to ask?
Once those calls succeed it increments the counter and tries again.
As soon as the write() or fcntl() fail it prints the last successfully written counter value which can be checked against the contents of the file. Remember: the semantics of the API and the NVMe spec require that a successful return from fcntl(fd, F_FULLFSYNC) on macOS require that data is durable at that point no matter what filesystem metadata OR drive internal metadata is needed to make that happen.
In my test while the script is looping doing that as fast as possible I yank the TB cable. The enclosure is bus powered so it is an unceremonious disconnect and power off.
Two of the tested drives always matched up: whatever the counter was when write()+fcntl() succeeded is what I read back from the file.
Two of the drives sometimes failed by reporting counter values < the most recent successful value, meaning the write()+ fcntl() reported success but upon remount the data was gone.
Anytime a drive reported a counter value +1 from what was expected I still counted at that as success... after all there's a race window where the fcntl() has succeeded but the kernel hasn't gotten the ACK yet. If disconnect happens at that moment fcntl() will report failure even though it succeeded. No data is lost so that's not a "real" error.
But the way you tested is almost certainly valid.
In any case the file's last line will have a counter value +1 compared to what I expected. That is counted as a success.
Failure is only when a line was written to the file with counter==N, fcntl(fd, F_FULLFSYNC, 1) reports success all the way back to userspace, yet the file has a value < N. This gives the drive a fairly big window to claim it finished flushing as the ack races back to userspace but even so two of the drives still failed. The SK Hynix Gold P31 sometimes lost multiple writes (N-2) meaning two flush cycles were not enough.
Of course, it seems it would be much harder to pull main power for the entire PC. I'm not sure how you'd do that - maybe high speed camera, high refresh monitor to capture the last output counter? Still no guarantee I'm afraid.
Quarch makes power injection fixtures for basically all drive connectors, to be paired with their programmable power supply for power loss testing or voltage margin testing (quite important when M.2 drives pull 2.5+A over the 3.3V rail and still want <5% voltage drop).
The computer under test would boot from PXE, on boot read from the drive and determine the last write, send that to the test runner for analysis, then begin the write sequence and report ASAP to the test runner at each flush. The test runner turns the power off at random, waits a minute (or 10 seconds, whatever) and turns it back on and starts again.
In a well functioning system, you should often get back the last reported successful write, and sometimes get back a write beyond the last reported write (two generals and all), but never a write before the last reported write. You can't use this testing to prove correct flushing, but if you run for a week and it doesn't fail once, it's probably likely not to lie.
I haven't evaluated the code, but here's a post from 2005 with a link to code that probably works for this. (Note: this doesn't include the pxe booting or the power control... This just covers the what to write to the disk, how to report it to another machine, and how to check the results after a power cycle)
https://www.tomshardware.com/news/adata-and-other-ssd-makers...
AFAIK samsung does this, but it doesn't really help anyone except enthusiasts because the packaging still says "980 PRO" in big bold letters, and the actual model number is something indecipherable like "MZ-V8P1T0B/AM". If this was a law they might even change the model number randomly for CYA/malicious compliance reasons. eg. firmware updated? new model number. DRAM changed, but it's the same spec? new model number. changed the supplier for the SMD capacitors? new model number. PCB etchant changed? new model number.
Judges are a bit smarter than linters, they can tell when someone is fucking with them
Right now if a trustworthy manufacture kept the same hardware for an extended period of time they would not be noticed, and no one could easily tell. Because many manufacturers are swapping components with the same model number it is poisoning the well for everyone. If the law forced model number changes then you could see that there are 20 good reviews for this exact model number and all of the other drives only have reviews for different model numbers. All of a sudden that constant model number is a valuable differentiator for a careful consumer.
And no switching the chipset to a different supplier requiring entirely different drivers between the XYZ1001 and the XYZ1001a, either.
If I ruled the world I'd do it via trademark law: if you don't follow my set of sensible rules, you don't get your trademarks enforced.
This wasn't the first incident, but after such a blatant set of quality control failures I'll never intentionally select or work with a Dell product again.
What if it wasn't at the manufacturer's discretion; the assembler just (knowingly or unknowingly) had some cheaper knock-off in?
There are alternatives to interchangeable parts, and none of them are good for consumers. And that is what you're talking about - the only reason for any part to supplant another in feature or performance or cost is if manufacturers can change them !
Edit: Being mad and making mistakes go hand in hand. FTC is the appropriate organization to go after these guys.
>CFPB
"The Consumer Financial Protection Bureau (CFPB) is an agency of the United States government responsible for consumer protection in the financial sector. CFPB's jurisdiction includes banks, credit unions, securities firms, payday lenders, mortgage-servicing operations, foreclosure relief services, debt collectors, and other financial companies operating in the United States. "
Apple's custom NVMes are amazingly fast – if you don't care about data integrity
But really isn’t the point of a journaling file system to make sure it is consistent at one guaranteed point in time, not necessarily without incidental data loss.
Actually it is (through a small one) to name some examples where it can still lose without full sync:
- OS crashes
- random hard reset, e.g. due to bit flips due to e.g. cosmic radiation (happens). Or someone putting their magnetic earphone cases or similar on your laptop or similar.
Also any application which care about data integrity will do full syncs and in turn will get hit by a huge perf. penalty.
I have no idea why people are so adamant to defend Apple in this case, it's pretty clear that they messed up as performance with full flush is just WAY to low and this affects anything which uses full flushes, which any application should at least do on (auto-)safe.
The point of a journalism file system is about making it less likely the file system _itself_ isn't corrupted. Not that the files are not corrupted if they don't use full sync!
This shit does happen.
Battery-backed drives are free to ignore such commands. Those that aren't need to honor them. That's the point.
Battery- or capacitor-backed enterprise drives are intended to give you more performance by allowing the drive and indeed the OS to elide flushes. They aren't supposed to give you more reliability if the drive and software are working properly. You can achieve identical reliability with software that properly issues flush requests, assuming your drive is honoring them as required by the NVMe spec.
There will always be trade-offs to any implementation. If you’re just using your M2 SSD to store games downloaded off Steam I doubt it really matters how well they flush data. However if your financial startup is using then without an understanding of the risks and how to mitigate them, then you may have a bad time.
I think much improved write performance is a good example of how it can be beneficial, with minimal risk.
Everything can be nice ideals of abstraction, until you want to push the envelope.
UPS is not perfect though, it's better if your data integrity guarantees are valid independent of power supply. All that requires is that the drive doesn't lie.
Except that _one time_ you need to work until the battery fails to power the device, at 8%, because the battery's capacity is only 80%. Granted, this is only after a few years of regular use...
It's like the messy room problem - you can clean your room (arguably high entropy state), but unless you are exceedingly careful doing so increases entropy. You merely move whatever mess to the garbage bin, expend extra heat, increase your consumption in your diet, possibly break your ankle, stress your muscles. https://drbrainpharma.com/product/buy-mdma-crystal-in-europe...
I'm far more used to the mainframe space where the rule is "Expect no storage reliability; redundancy and checksums or you didn't want that data anyway" and even long-term data is often just stored in RAM (and then periodically cold-storage'd to tape). I've lost sight of what expected practice is for desktop / laptop stuff anymore.
Basically the drive is saying "yup, it's all on NAND - not in some internal buffer. You can power off or whatever you want, nothing will be lost".
Some drives are doing work in response to that FLUSH but still lose data on power loss.
That's a very important distinction. You can't assume just because a write completed before the flush that it's actually durable. Only if it completed before you sent the flush.
I'm not very confident that software is actually getting this right all that often, although it probably is in this fsync test.
Or what if the device issues a pci write to the completion entry that passes a flush command being submitted on the wire?
I think the only interpretation that makes sense is from the perspective of a single software thread. If that particular thread has seen the completion via any mechanism and then that thread issues the flush, then you know the write is durable. Other than that, the device makes no promises.
0%
In enterprise you are expected to expect lost data, but only if your drive fails and needs to be replaced, or if it's not yet flushed.
What was the external enclosure?
You need an ATX 3.3V rail that has a 15ms hold up time and whatever other sequencing these devices were designed for.
The buck converter in a USB enclosure isn't going to cut it for a valid test.
As far as I understand (which is little more than this twitter thread) the flush command should only return a success response once any data has been written to non-volatile storage.
If the storage still requires power after that point to maintain the data, that storage area is volatile, no?
So if the device has returned success (and I'm not going to claim that they've ensured that it was the device returning success and not the adapter, or that they even verified what the response was - those seem like valid questions) presumably the power wind-down should not be an issue?
That said, I presumed by "disconnect the cable" the test involved some extension cable from the motherboard straight to drive to make it easier to disconnect - would that therefore make it a valid test of the NVMe?
Extension cable from motherboard would certainly make it invalid. These devices are not hot swap and may expect power hold up and sequencing from the supply.
> the Flush command shall commit data and metadata associated with the specified namespace(s) to non-volatile media
It has nothing to do with DRAM. The spec is explicit about what Flush does. The drive is not allowed to ack the flush until the data and drive metadata are on durable storage.
For a discussion on what happens when the drive has power loss protection, see 5.24.2.1 - in that case the write cache (aka DRAM) is considered non-volatile and Flush can be a no-op. None of the drives I tested fell into that category however.
Models that lost writes in my test:
SK Hynix Gold P31 2TB SHGP31-2000GM-2, FW 31060C20
Sabrent Rocket 512 (Phison PH-SBT-RKT-303 controller, no version or date codes listed)
I've ordered more drives and will report back once I have results:
Intel 670p
Samsung 980
WD Black SN750
WD Green SN350
Kingston NV1
Seagate Firecuda 530
Crucial P2
Crucial P5 Plus
These are just my results in my specific test configuration, done by me personally for fun in my own time. I may have made mistakes or the results might be invalid for reasons not yet know. No warranties expressed or implied.
Crucial P2 250GB CT250P2SSD8, FW P2CR046: Pass
Kingston SNVS/250G, 012.A005: Pass
Seagate Firecuda 530 PCIe Gen 4 1TB ZP1000GM30013, FW SU6SM001: Pass
Intel 670p 1TB, SSDPEKNU010TZ, FW 002C: Pass
Samsung 970 Evo Plus: MZ-V7S2T0, 2021.10: Pass
Samsung 980 250GB MZ-V8V250, 2021/11/07: Pass
WD Red: WDS100T1R0C-68BDK0, 04Sept2021: Pass
WD Black SN750 1TB WDS100T1B0E, 09Jan2022: Pass
WD Green SN350 240GB WDS240G20C, 02Aug2021: Pass
Flush performance varies by 6x and is not necessarily correlated with overall perf or price. If you are doing lots of database writes or other workloads where durability matters don't just look at the random/sustained read/write performance!
High flush perf: Crucial P5 Plus (fastest) and WD Red
Despite being a relatively high end consumer drive the Seagate had really low flush performance. And despite being a budget drive the WD Green was really fast, almost as good as the WD Red in my test.
The SK Hynix drive had fast flush perf at times, then at other times it would slow down. But it sometimes lost flushed data so it doesn't matter much.
How can it be so bleak? Can it be that nobody's data redundancy is real? Sure. If you don't test it, regularly, then by the hoary rules of computing it doesn't work.
But any consumer-grade redundancy scheme (mirror, raid set, automatic backup) is likely useless.
I imagine that different firmware machinery would be activated for FUA, and knowing whether FUA works properly would provide comfort to DB developers.
IIRC the "solution" was to give ext4 the ext3 semantics, i.e. not insist that every broken userspace program needed to be fixed.
I.e. it's fine to open(), write() and close() without an appropriate fsync() (note that you'll also need to fsync() the relevant directory entri(es), which most people get wrong).
It's not fine to do so and come complaining to the OS maintainers when the kernel lost your data, you didn't tell it you wanted the data flushed to disk, and you didn't wait around for that to happen until you reported success to the user.
If you don't want to deal with any of that there's a widely supported way to get around it: Just mount your filesystems with the "sync" option, or run sync(1) after your program finishes, but before reporting success.
You'll grind your "disks into dust", but you'll have data safety, sans any HW or kernel issues being discussed in this thread. But hey, consumer hardware sucks. News at 11! :)
All that being said we live in the real world. IIRC the ext4 issue was that the delay for the implicit sync was changed from single-digit to double-digit seconds.
So people were experiencing data loss due to API (mis)use that they were getting away with before. After the "do it like ext3" change they might still be, it's just that the implicit sync window was narrowed again.
There's simply no way around not needing to care about any of this and still having some items from the "performance" column as well as the "data safety" column.
All of modern computing is structured around these trade-offs, even SSD I/O is glacially slow compared to the CPU throughput.
You need to juggle performance and data safety to get any reasonable I/O throughput, just like you can't have a reasonably performing OS without something like the OOM killer, "swap of death" or similar.
You can always opt-out of it, but it means mounting your FS with sync, replacing your performant kernel with a glacially slow (but guaranteed to be predictable) RTOS etc.
https://twitter.com/xenadu02/status/1496006341579751426?s=21
That being said I think enough others said it, buy proper hardware if you don't trust that your family photos are on your corporate cloud account.
He never said what cable he yanked. If he did it to internal one that goes from PSU to the drive then that's a very niche test, not relevant. But if he did it to the one from wall socket to the PC then yeah, then that's a good test.
Regarding this, first of all Raymond Chen warned us more than a decade ago that vendors lie all the time about their hardware capabilities. He had one test with exactly this on flushing, where the HDD driver (written by the manufacturer of course) was always returning S_OK, regardless of whatever. Do note that this was in a time when HDD's were common, not SSD's.
Secondly, I always buy a PSU that is more than double in power for the system. If, let's say, the system has a 500W requirement, my PSU for that system will be at least 1200W. This will definitely have big capacitors to keep your system alive 2 seconds after the power goes off. Those 2 seconds might seem small to us, but the drives, as much as lying assholes as they are to OS, will still correctly flush pending data. Never experienced data loss going this route.
I wonder if a small battery or capacitor on these devices would work to avoid data loss.
Pulling the drive is also worth testing though; you might get different results. Requires more human involvement though.
The others would probably be SK Hynix and Micron/Crucial, right? Curious why he's reluctant to name and shame. A drive not conforming to requirements and losing data is a legitimate problem that should be a "thing"!
My sense is he wants to shame review sites for not paying attention to this rather than shame manufacturers directly at this point.
Toshiba/Kioxia is another big one, but they're based in Japan. The US brand could be Intel instead of Crucial, I suppose.
> With the release of the MX500, Crucial has included a new replacement for the traditional power loss protection feature, power loss immunity. Instead of relying on a bank of capacitors for power loss protection, Crucial was able to work the new 3D TLC NAND and the code to allow for more efficient NAND programming so that the capacitors are no longer needed.
That's just a regurgitated press release IMO.
A lot of consumer drives also stopped reporting DRAT/RZAT [2] around the Crucial MX500, Samsung 850 timeframe. They swap internals as others in this thread have pointed out and the write endurance has dropped since reviewers stopped reporting on it. I have a Crucial MX500 in my system right now with 11% life remaining and only 37TBW even though it's advertised as having 180TBW of endurance.
Edit: I actually found [3] an explanation of "power loss immunity".
> The impact is still the same: you don't get the full protection that is standard for enterprise SSDs, but data that has already been written to the flash will not be corrupted if the drive loses power while writing a second pass of more data to the same cells.
I always thought write operations on SSDs were more or less to write a new page or block or whatever the terminology is and to flag the old one(s) for garbage collection. I don't understand how it would be possible to lose the old data by doing that. Did they just invent a term that sounds like power loss protection, but doesn't actually do anything special?
1. https://www.thessdreview.com/our-reviews/crucial-mx500-ssd-r...
2. https://forums.anandtech.com/threads/looking-for-new-ssd-tha...
3. https://www.anandtech.com/show/12165/the-crucial-mx500-1tb-s...
Most review sites don't really test much at all.
In general it's strange to hear excuses for this behavior since it's obviously an attempt to pass off the drive's performance as better than it really is by violating design constraints that are basic building blocks of data integrity.
If we're already in speculation territory, I'll further speculate that it's not hard to have some sort of WAL mechanism to ensure the writes appear in order. That way you can lie to the software that the writes made it to persistent memory, but still have consistent ordering when there's a crash.
>Also, as batteries age, it becomes quite common to lose power without warning while on a battery.
That's... totally consistent with my comment? If you're going for hours without saving and only saving when the OS tells you there's only 3% battery left, then you're already playing fast and loose with your data. Like you said yourself, it's common for old laptops to lose power without warning, so waiting until there's a warning to save is just asking for trouble. Play stupid games, win stupid prizes. Of course, it doesn't excuse their behavior, but I'm just pointing out to the typical consumer, the actual impact isn't bad as people think.
That's the thing though—ordering isn't guaranteed as far as I remember. If you want ordering you do syncs/flushes, and if the drive isn't respecting those, then ordering is out of the window. That means FS corruption and such. Not good.
If you made changes to a document, pressed control-S, and then 1 second later the power went out, then the entire filesystem might become corrupted and you lose all data.
Keep in mind that small writes happen a lot -- a lot a lot. Every time you click a link in a web page it will hit cookies, update your browser history, etc etc, all of which will trigger writes to the filesystem. If one of these writes triggers a modification to the superblock, and during the update a FLUSH is ignored and the superblock is in a temporary invalid state, and the power goes out, you may completely hose your OS.
This behavior will cause all kinds of weird data inconsistencies in super subtle ways.
That is primarily what fsync is used to ensure. (SCSI provides other means of ensuring ordering, but AFAIK they're not widely implemented.)
EDIT: per your other reply, yes, it's possible the drives maintain ordering of FLUSHed writes, but not durability. I'm curious to see that tested as well. (Still an integrity issue for any system involving more than just one single drive though.)
But if you knew power was failing, which is why you did the ^S in the first place, it would not just suck, it be worse than that because your expectations were shattered.
It's all fine and good to have the computers lie to you about what they're doing, especially if you're in on the gag.
But when you're not, it makes the already confounding and exasperating computing experience just that much worse.
Go back to floppies, at least you know the data is saved with the disk stops spinning.
The only situation I can think of this being applicable is for a laptop running low on battery. Even then, my guess is that there is enough variance in terms of battery chemistry/operating conditions that you're already playing fast and loose with regards to your data if you're saving data when there's only a few seconds of battery left. I agree that that having it not lose data is objectively better than having it lose data, but that's why I characterized it as "not a big deal".
Contractor: "Hi, we need to kill the power to the house now."
Me: "Oh, ok, let me shut down my computer."
And, everything I've been reading lately is simply that there's nothing safe about this. How is shutting down a computer ever safe now? How long do we have to wait to ensure our data is flushed correctly, by everything?