How Reliable Are SSDs?
backblaze.com
backblaze.com
This article was clickbait. You expect their typical blog post quality, and instead get "An SSD should last as long as its manufacturer expects it to last".
Hope they step it up for the next post (or are not afraid to stay quiet until they have some interesting data to write about)
I can echo that, since reading those articles, I have been using B2 more, and gifting years of BackBlaze to beloved lusers. I think it just takes time to generate good data like this though, and it seems that SSDs from the last ~3 years have considerably lower rates of unforeseen failure than hard disks, which makes measuring failures more difficult.
Then again there was a huge difference in failure between people who do SSD well (intel and sandisk) and manufacturers who did things... less well. I remember the OCZ controller would crash and block write operations for a half second while it rebooted. Kinda messed with the write throughout numbers.
There's something majestic about a datacenter with no moving parts except fans.
Step 1: Assume every home user doesn't use backup. Most users have at least one important thing on their hard drive that is not getting backed up right now, either because they don't have a backup solution, or because the thing is located in some non-standard location that is not getting backed up. I can think of one thing on my own computer that qualifies for the latter case right now.
Step 2: Now write the article taking step 1 into account. Assume somebody reading this article is going to use it to decide where to store their unbacked up important data.
Nobody cares how long the average case is when its longer than they plan to use the drive.
They care about the likelihood their unbacked up data will be lost because their ssd died before its time.
They care about how well the firmware is designed to handle error cases, does it shut down and refuse to power up or does it go read only? Does minor localized errors in some data cause it to lock out access to all data?
Those are the interesting questions wrt to ssds.
If everybody shares that opinion, no wonder software is getting more and more bloated, slow, and unresponsive. My main work laptop is an i3 from 2010. I just replace the hard drives and battery every couple of years. Why on earth should I change the whole machine each 3 years? If everyone did that, that would be terribly wasteful.
What is the primary cause of SSD failures?
Is it flash wear out of all cells, leaving no good cells to write new data as people talk about?
Is it flash wear out only of some critical cells?
Is it that cells degrade over time while the drive is off, and when next powered up, they are too far gone to recover the data?
Is it unrecoverable corruption of critical internal data structures?
Is it unrecoverable corruption of user data (IE. The user could reformat the drive and have it back to a usable state)
Is it hardware failure outside the typical 'flash wear out' model?
Is it firmware bugs (eg. A badly timed power off leaves data in such a state that the firmware can't initialize next time)
Throw some nvme and optane in there, see if we can wear one of those out.
So now we do SSD in front of entire disk systems, do we get the same benefit? Does a rust RAID run more reliabily if an SSD acts as frontline storage in some write through model?
https://www.circonus.com/2019/01/which-block-i-o-scheduler-i...