SSD Failures in Datacenters
dl.acm.org
dl.acm.org
The HP RAID controller failed to detect corrupted writes on what I seem to remember were Intel SSDs. The only way we learned about failed SSD drives is when we started catching large numbers of DB page checksum errors, and by that time it was too late to do anything about it.
Operationally it sucked, swapping out drives is a lot easier maintenance than having to rebuild a DB from backup.
(I mean this as an honest question, I'm interested to know if there are performance or reliability gains to be had from a hardware controller)
Also, these machines were built by DB people, who have traditionally placed a lot of value in hardware RAID controllers, and to be fair, the HP controllers worked well with HDDs. Also ZFS wasn't a realistic option for production Linux in 2013. Me personally, I'm not a fan of exotic proprietary hardware, I much prefer working with JBOD disk setups and software RAID where appropriate.
Don't all enterprise SSDs come with their own backup capacitors?
Yes, but the cost of such drives makes rolling the dice with consumer grade drives appealing to some[1].
From the same thread, here's a word of warning from a Stackoverflow devops engineer[2] wrt to lack of capicitors on consumer grade drives and the perfect storm that can occur when UPS goes down in a power failure (unlike with spinning disks hardware RAID is unable prevent data corruption/loss in this case).
[1] https://laur.ie/blog/2015/06/ssds-a-gift-and-a-curse/
[2] https://laur.ie/blog/2015/06/ssds-a-gift-and-a-curse/#commen...
I don't think that professional devices are 100% fail-proof either. I vaguely remember some test where they setup a test with forced power failures during heavy writes, and they zapped all drives but a few eventually.
Something that worried me was that the failure mode quite often seemed to be a totally bricked SSD.
The fine-print is important here. Crucial (IIRC) advertises some of its low-end consumer series with power-loss protection… against internal data corruption of the Flash Translation Layer. AKA "We won't completely brick the entire device if you pull power!" Err… thank you? Why are you advertising this as feature?
What is the point of all that complication if the block device is just going to throw away those guarantees? If we're giving up on power loss safety I can make a super fast filesystem by making fsync a no-op and keeping everything in ram until I feel like writing it.
Same with HDDs. That's why battery-backed hardware RAIDs (+ disabling this cache on the connected disks) are a thing in the first place.
> What is the point of all that complication if the block device is just going to throw away those guarantees?
Ask that storage manufacturers.
Or would this functionality somehow be counteracted by a cache inside the SSD?
Perhaps filesystems should include a "ssd" mount option, because there might be more things to worry about (for example frequent writes wearing out the device).
They do. You can, for instance, mount a UFS filesystem "sync", which means that all writes are synchronous. I assume extX has a similar option ...
Also be careful with sync on SSD media; this is what the mount manpage says about the sync option:
> In case of media with limited number of write cycles (e.g. some flash drives) "sync" may cause life-cycle shortening.
https://www.usenix.org/conference/fast16/technical-sessions/...
Reference #50.d66d717.1467850467.45a102d"
I supposed their ssd failed
:-)
tl:dr the benefits were astounding. The failure rates were super low. We saved something like 2 million dollars by significantly increasing the I/O performance of the fleet preventing the need to re-shard.
Choosing a drive We went with an Intel MLC (consumer class) drives. No other drive had such low DOA rates and good match between performance and price. We set the max lba to 80% of available capacity. (Actually recommended by Intel) This change eliminated the pathological case where a full drive continually overwrites the same sectors. The consumer drives suddenly had a lifespan comparable to enterprise drives (though a slightly slower read speed- which was fine because we needed balanced reads and writes) We also exported and monitored the wear leveling stats (available via a SMART value) so we won't run into the case where they all unexpectedly wear down at the same time. The projected lifespan looked really good and in most cases exceeded that of the servers. (in retrospect this turned out to be true)
Raid config We ran in Raid 0 (because I love to live on the edge- No, actually because we split the boot, data and bin log volumes and data integrity was preserved by redundancy within the the fleet.) The conversation I had with the dbas about this was one of the most eye opening conversations of my career. Turns out if a replica failed or lagged too much they simply tossed the data and started a new one. They didn't actually need or want RAID10 and that changes the economics of the picture significantly.
This would have been 4 years ago now. We deployed several thousand drives. Of that only one ended up with abnormally high wear level. (I proactively threw it out when we decommissioned the hardware to prevent a problem for the next guy). I had a handful of drives which didn't pass the pre-service checks and a another handful that failed in production.
Compare that to 10k spinners which appear to have an annual failure rate of 1 in 20.
SSD Failure profile: We would typically see a drive failing to respond properly (announcing a crazy size, not showing up at all, failing to save its max LBA). Most of this happened in the pre-production check phase. Post deployment failures you can count on one hand (I have the normal number of fingers ;-)
Qual process We qualified 6-8 different manufacturers, high end enterprise to the dirt cheap vendor with a three letter name now synonymous with garbage (I won't even say the name here). The cheap drive had high DOA rates, and the unsettling tendency to spontaneously reset itself causing a several second pause (not so good for our uses, but might be tolerable in other circumstances)
6. Concluding Remarks This paper presents an extensive characterization of SSD failures using field data.We first show that SSD failure rates in the field can be very different from what vendors specify. Next, we identify and quantify four types of SMART failure symptoms exhibited by the SSDs, and provide characteristics of symptom occurrence, intensity and progression rate. We show that despite their presence, the symptoms alone cannot be a sufficient indicator of failures. We have also studied the impact of multiple provisioning and operational factors across different layers of the datacenter hierarchy on SSD reliability. Many of these factors are individually influencing, and can also interact with each other in complicated ways, to impact SSD failures. We have used machine learning and graphical model based approaches to systematically consider the impact of multiple influential factors towards answering the what, when and why of SSD failures. We believe the insights gained from this paper can greatly influence the design, provisioning and operational decisions for SSDs in datacenters.
Failure rates seems to not fit the models and be a tad higher than expected.
Most SSD are reliable, the only problem is they are failing in a way that is tough to detect leading to consequences we may ignore. (falsely positive functioning SSD in operations and we may experience silent corruptions)
Given the "too many knobs" of the SSD they can defect in non predictable snowballing effects (chaotic and dramatic).
Basically this study confirms that IEEE standards for rating MTBF are to be reconsidered drastically and that our models of failure for SSD are far from totally being well understood and that SSD is in production while all costs are not totally yet known. The Cost of SSD specific failures is not yet known and they urge SSD makers to begin studying the effects of their "knobs" on reliability
Even though it is written in a very scientific neutral tone trying not to scare people, it basically can be used as a strong evidence to ban SSD from critical systems.
It basically is an heavy blow to SSD industry since it attacks it on the costs model that is based on its expected reliability and says that these drives are basically still an unknown territory when it comes to failures.
If SSDs make your business 2x as productive or 2x as remunerative, the cost of an extra backup or two is simply a blip.
> ban SSD from critical systems.
That's an over-reaction if I've ever seen one. I guess it depends what a "critical system" is for you. If it's something that has to guarantee high performance under massive load, SSDs are what you need, period. If it's something that has to guarantee 100-year data storage under a mountain, then there might be better choices, but they might not involve spinning drives either.
Ok, it may be RAID1+0 but, what about data corruptions?
If primary fails, and secondary is "ok" what will be the state of your data? You don't roll back an image when everything is green.
The point is that in critical systems according to the SLA or the critical aspect of the mission (flight recorder, medical instruments...) silent failures, silent data corruption may not be ok.
However for 99% of the business that is making money out of SSD right now in the world (advertisement, marketing) this is not really a concern. They don't aim at an absolute reliability. Most industry just aim at good enough.
Some small islands of IT aims at doing a legally accurate book-keeping. Like banking, accounting, legal electronic archiving for instance...
This topic is covered in Signal Processing in the first chapter of the cost of False Positive and False Negative.
Security and risk management norms force to enumerate and evaluate the costs and predicting the outcomes of failures in terms of costs.
Here, the paper is saying SSD works pretty much fine all the time. But for failure, there is a lot we don't know.
How do you evaluate the cost of a risk you cannot compute?
And how does another solution that maybe slower rates in risk management with a better failure model that can be better predicted?
The scarriest part in the study is the implication that their could be propagation of failures by coupling in certain environment and that some failure could be snowballing.
SSD are affected by heat/memory usage, CPU load is affected by failures, heat rises, SSD are more likely to fail ....
They are implying that failures can propagate or spread or happen like a contamination.
1 SSD failing alone is already a problem, even small chances of coupling in failures is making the potentiality for a massive failure.
Would you take that risk with a SLA of 99.999%?
As I said, it depends on what your definition of "critical system" is. For aviation, for instance, there are many other factors (shielding etc); no hardware is perfect for all uses.
> How do you evaluate the cost of a risk you cannot compute?
Easy: you come up with a worst-case scenario, make plans to address it, and the cost of implementing such plans is your cost. Say you assume you could get silent errors: you build test routines for actual data integrity and implement test systems for continuous checks. The cost of building and maintaining all that is the cost of (accepting) your risk.
Point-in-time recovery is a snap compared to that domain.
Aviation standard goal is 9 nines of reliability, and many systems get there.
> > ban SSD from critical systems.
> That's an over-reaction if I've ever seen one.
No, it's really not. SFJulie's point is that we don't know enough about the way SSDs fail to build a good failure model for them, and that we currently don't seem to have a good way to even detect that they are failing. If you can't build a failure model for something, you can't build a good risk model including it. That's kinda her point.
FWIW, I don't know that I fully agree with that point: it seems to me that cryptographic integrity checks and backups would suffice, but … maybe not?
> If it's something that has to guarantee high performance under massive load, SSDs are what you need, period.
We seemed to do alright four years ago, no? SSDs are not a requirement: one can get equivalent system performance, with predictable failures, by spending more money on more hard drives and CPUs to talk to them.
To me the three scariest words in engineering are "common mode failure". I personally never buy more than one hard disk at a time, so that I'll hopefully avoid something like both haven been shaken too much in their shipping container (the best explanation anyone was able to come up with for a Seagate failure pattern some time ago).
So, yes, the above, to detect and recover, but plenty of services need very up to date backups, so if I was designing such a system, I'd make sure the (near-)real time backup would be to SSDs that share nothing in common with my production ones, different manufacturer, chip sets, etc.
Or if the load could handle it ... ah, I see why Seagate is still selling real SAS enterprise disks (last time I checked, if I remember correctly, <= 1TB and up to 15K RPM). We understand those, not that they won't also fail on you.
Basically all you need to define is the expected output. "99.9pct of my reads are within 1ms" + "The read response is correct 99.9999% of the time" and then verify. Eject the disk if that doesn't hold, restore from replica. The rate of replacements is a factor of how close you are to the true failure model.
Lots more money. Many, many more hard drives, with failure rates to match. And you're still not going to get there.
We started running a 240TB all SSD array a few months ago and we get sub-millisecond latencies on most reads and writes. When we see an excursion to 10ms that's cause for concern; 50ms and we do a root-cause. We can peg multiple 16GBit fiber channels and it's okay. And it fits in half a rack; try that with disk.
I cannot imagine doing this with spinning rust.
These SSDs are not failing for physical reasons, which is the model we have in our minds for HDD failure.
They are failing for logical reasons.
In the near term, I think there is a very simple way to protect against this and that is to mix models. At any given time, you can buy the very latest Intel SSD and also the one-model-ago Intel SSD on Amazon. They're both fantastic and wonderfully performant and their sizes are typically identical.
So, for instance, when we make a boot mirror out of SSDs, we buy one of each and make our mirror with those two.
Now we know that if there is some silent, insidious, bizarro firmware bug that makes the drive brick or causes it to burn itself out early, it can't affect them both identically (which is important, since they have identical workloads, since they are mirrors).
This isn't as useful on a large array since there are not 12 or 15 reliable manufacturers to mix between - but in a large array, you aren't giving each drive an identical workload like you are with a mirror.
I'll bet a lot of SSD usage is confined to boot mirrors, so I think this is a decent practice that a lot of people can adopt.
http://delivery.acm.org/10.1145/2930000/2928278/a7-narayanan...