There are at least these two reasons: proper testing of endurance of a device is not possible, you can only test that in a pretend kind of way. We are talking about years of service that have to be somehow emulated in at least weeks... and, of course, you cannot really account for physical behavior of materials from which the disk is made by simply running multiple I/O workloads. Now, add to this that the larger the storage capacity, the harder it is to check durability because throughput becomes the bottleneck. I.e. if you emulate device wear by running more I/O workloads, then, proportionately to the size of the device, you will be able to run fewer of them per unit of storage, because you are bounded by the throughput.
Any real (good) tests so far rely on previous generation of devices, and don't necessarily reflect the current situation. I.e. if your SSDs survived for only five years, who's to say that the next generation you will buy will last more or less? They are very likely not the same kind at all...
Also, it's really silly to measure disk durability in units of time... I mean, most tests that intend to measure durability model it by running I/O workloads, so, they'd typically measure durability in something like "how many times can a unit of storage be written over". The usage patterns vary dramatically across different kinds of workloads. So, if, eg, you are running a build server, you will wear your storage a lot faster than if you are running a (well-configured) database server, and still much faster than if you were running a video streaming service, and still much faster if that's a (well-configured) Web server... and the difference could be an order of magnitude between these.
If it is not plugged in, i.e. unpowered, don’t count on anything more than three months. Thread with sources: <https://news.ycombinator.com/item?id=27573332>
While SSDs are inferior for long term data persistence if it’s irreplaceable data you should be bring the file system online on a schedule and let it run checksums. If something errors restore from another copy. The magic of digital storage is perfect copies at low cost. The loss of any one copy is fine as long as it has been copied forward elsewhere.
"Remember that the figures presented here are for a drive that has already passed its endurance rating, so for new drives the data retention is considerably higher, typically over ten years for MLC NAND based SSDs."
It is kinda obvious that if you pass the numbers considered safe then bad things shall happen.
But a SSD sitting on a shelf starting to lose bits after three months seems incredibly low.
People don't know about that and they'd be screaming by now.
I don't think it's that bad.
In that thread is also this post:
<https://news.ycombinator.com/item?id=27573332#27573720>
Which contains this link:
<https://web.archive.org/web/20210502042514/http://www.dell.c...>
Which in turn contains this text and table:
I have unplugged my SSD drive and put it into storage. How long can I expect the drive to retain my data without needing to plug the drive back in?
It depends on the how much the flash has been used (P/E cycle used), type of flash, and storage temperature. In MLC and SLC, this can be as low as 3 months and best case can be more than 10 years. The retention is highly dependent on temperature and workload.
┌────────────────┬─────────────────────────────────┐
│NAND Technology │ Data Retention @ rated P/E cycle│
├────────────────┼─────────────────────────────────┤
│SLC │ 6 Months │
├────────────────┼─────────────────────────────────┤
│eMLC │ 3 months │
├────────────────┼─────────────────────────────────┤
│MLC │ 3 Months │
└────────────────┴─────────────────────────────────┘Modern hard drives have warrantee limits on how many bits you can read/write during their service lifetime, since they lower the head from about 10nm to 1-2nm during reads and writes, and head lifetime is correlated with the number of hours it spends at that 1-2nm height.
With workload specifications in the range of 500TB/year (IIRC - I took a quick look and couldn't find any recent specs) that works out to less than QLC levels of endurance. It's not the same, though - if you read/write every byte of a hard drive 300(?) times the failure rate rises and the vendor gets nervous; if you overwrite QLC 3000 times it's on its last legs, and 1 or 2K more writes will almost certainly kill it.
SSDs have quite predictable durability, mostly because the unit that fails is less than 1GB, so the law of large numbers kicks in.
Note also that "durability" is a soft target - the normal failure mode is that blocks retain their data for shorter and shorter times before hitting the ECC error limit, so if your storage system moves data around every few months you can push the flash harder than in e.g. a laptop, where you don't want to risk losing all your data if it sits powered down on a shelf for half a year.
You're forgetting the controller, which has no qualms with dying 100% unexpectedly. I don't know why consumer SSDs have such poor quality controllers that they can randomly die, something that HDDs seemingly haven't struggled with in decades.
Without access to vendor tools (like a JTAG debugger and source for the controller) I'm not sure it's possible to tell whether the controller itself failed, or it just decided not to wake up because the flash was dead.
But yeah, it sucks.
Finally, I'd note that a lot of earlier SSD vendors came out of the USB device market, where they were used to making things that had the reliability requirements of your average Happy Meal toy. There are only 2.5 hard drive vendors or so at present, in large part because the ones who weren't fairly good at reliability are dead now.