Analysis of SSD Reliability during power outages
lkcl.net
lkcl.net
Source: http://hardware.slashdot.org/story/13/12/27/208249/power-los...
I have my doubts that many of them would survive thousands of power cycles per day.
It's the SSDs that I worry about - some store their own firmware on the same NAND as the User's data, and allow it to become corrupted during power-loss events. Several cheap models won't last more than a couple days before they simply drop off the SATA bus and never come back.
The test is pretty pointless.
Instead of buying thousands of disks to test this scenario, one reasonable shortcut might be to repeatedly test a single disk.
Even a spinning disk can handle a few thousand power cycles in its lifetime, surely?
What this guy did was test the drives in a manner they weren't designed for. Sure, I can drop Corvettes off a 20 story building, then bitch about the results, but that wouldn't change the fact that my test was flawed from the onset.
All he did was subject something to an environment is wasn't designed for.
Sure, I can drop a Corvette from a 20 story building, but there's nothing to be gained when the crumple zone is packed into the tail lights.
The claim was that an SSD somehow converted more electric energy into heat immediately after power-up which would damage the SSD, so real consumption, not just a current peak that goes into storage for later consumption. Normal-ESR electrolytics might have a heat problem when used at a few kHz in switching applications, but certainly not at 0.1 Hz.
Executive summary: powerup stress is not an issue unless the drive was designed by a moron.
Essentially, it sounds like they have a single unit that gets turned on, gathers sensor readings, and then get turned off. I'm guessing these are not handheld devices, otherwise they most likely would have a battery attached. If a non 24/7 factory, these could be turned off with the rest of the assembly line by someone throwing breakers off on the way out (very common).
- OCZ Vertex 4 SSD does (but this is not advertised)
- OCZ Vertex 3 does not, Vertex 3 Pro does
- Ditto for Vertex 2 / Vertex 2 Pro
- OCZ Deneva(2) R does, other OCZ Deneva/Deneva 2 do not -- and in the case of Deneva 2 R this is advertized
- Intel 320 does (but is only sata II), but this is not advertised at all
- Intel 520/530 does not
- Intel 330/335: unclear
- Intel S3700/S3500: does and this is advertised
Background: I was looking for an SSD to use for a ZFS intent log ("ZIL" -- ZFS's write ahead log) -- my requirements were a sandforce controller (or equivalent) and toggle NAND (so that I could use the same disk for both ZFS cache and the ZIL), and a supercap. This was surprisingly hard to find.
What I'd like to see is:
1) Data on which SSDs have supercap
2) Data on which SSDs actually honour cache flush request (then a UPS + forcing a cache flush request upon power failure + redundant power supplies/multiple replicas in a distributed system) would suffice.
3) Best yet: have an API for for check whether or not an SSD has a supercap, whether or not the supercap still holds charge, and policy for honouring cache flush requests. Let the OS decide based on a policy I set.
If you are a consumer, the practical recommendation is not very different from the practical recommendation I'd give to anyone using spinning disks: RAID1/1+0/RAIDZ(-2) (with SSDs coming from different patches so that they do not wear out at the same time), UPS (for power outages), backups (against yourself and against power supply failures).
For production: obviously use a UPS, put the WAL of your database on an SSD with a supercap, make sure that your database fsync()'s the WAL at a reasonable interval (on every transaction is probably unreasonable, but so is once an hour), and use a distributed system that replicates the WAL. If using a distributed system with (semi-)synchronous WAL replication is not an option and losing incremental data is not acceptable, use redundant power supplies.
All of the M500 line includes hardware AES 256-bit encryption, and Micron showed us an array of small capacitors on one of the M.2 form factor drives that supported flushing of all data to the NAND in the event of a power loss--not a super capacitor as seen in enterprise class SSDs, but there's no RAM cache to flush so it's just an extra precaution to ensure all of the data writes complete. http://www.anandtech.com/show/6614/microncrucial-announces-m...
Other features that set the M500 apart also center on optimizations that are clearly holdovers from the enterprise version of the SSD. Power hold-up is provided by a small row of capacitors that will flush all data in transit to the NAND in the event of a power loss. This feature is not standard with any other consumer SSD on the market, and in enterprise SSDs power capacitors typically command a much higher price structure. Finding power loss protection on the consumer M500 is a nice surprise, and one that users will need more often than they think. http://www.hardocp.com/article/2013/05/28/crucial_m500_480gb...
If you're talking about PostgreSQL, the general consensus from the mailing lists and IRC discussions seems to be that you should in fact not put your WAL on SSDs. $PGDATA/base, yes; $PGDATA/pg_xlog, no.
The amount of write churn in your WAL will burn through your SSDs wear leveling very quickly, and given that WAL is nearly perfectly sequential in access pattern, you lose (forfeit) nearly all the random IO benefit an SSD buys you.
My production master/slave pair only have SSD in them, so I'm "doing this wrong", given that advice. Our SSDs are SLC NAND, however, so their write endurance is vastly, vastly higher than the drives under consideration in both your comment, and TFA.
What I'd like to see is:
1) Data on which SSDs have supercap
Every SSD review I've read at Anandtech.com or StorageReview.com has mentioned the presence of a supercap on drives that have them.When they review consumer-oriented drives without the supercap, they typically don't say "there's no supercap" but 1) there are usually physical teardown pics so you can fairly easily check for yourself 2) be pretty sure there isn't one if they don't mention it, because they always talk about it if there is one.
This doesn't help with this, of course:
2) Data on which SSDs actually honour cache flush request (then a UPS + forcing a cache flush request upon power failure + redundant power supplies/multiple replicas in a distributed system) would suffice.
Would love to see reviewers test that as well. For now it seems like there's no way to know unless you're willing to do your own power-cycle torture testing.Also be sure to plug each redundant power supply into a separate UPS, and avoid equipment that has redundant power supplies on a single feed. UPS failures are a bear.
I have always wondered why Hosting companies aren't choosing Intel DC Series at all. As far as i know, one of the few hosting company using it are Hivelocity.
Of course, a RAID is only useful when the failure is drastic enough for the controller to notice it, which some of these power-down failures may not be.
I'm sure there's some SMART flag that shows it. No idea which though, not one of them is properly documented.
Spoiler: He only tested these five drives, only Intel survived, so if they are your candidates, apply his conclusion:
Crucial M4
Toshiba THNSNH060GCS 60gb
Innodisk 3MP Sata Slim
OCZ Vertex 32gb
Intel 320 and S3500
Notably missing is the Intel runner-up Samsung and probably others i'm not aware of, as well as other models.Did I miss an appropriate Samsung drive in my quick searches? Or is there reason to suspect that a Samsung drive that doesn't claim to have power loss protection would nonetheless handle this case better than the non-Intel drives that did make that claim? Because if not, then I don't think not evaluating Samsung drives compromises the results in any way.
So yeah; it's unfortunately bogus science - he's drawing invalid conclusions from a far too small sample.
One of the restrictions is price, which makes little sense in concluding so strongly for a measure involving quality. Not to mention that only a single specimen of a single model of each brand was tested. The top-voted comment in the thread points out some Intel models that don't have supercapacitors in them.
Already the article is updated with a couple of other drives to test. I guess that wasn't "End of discussion" after all...
He tested a bunch of equipment, and was kind enough to share his methods and results. Lot of people/companies don't do that.
He's not misrepresenting what he did, and he provided some valuable data, so kudos to him.
The worst case, which I've experienced with direct writing to flash chips, is that a totally unexpected flash location is corrupted. My guess was that the CPU started the write cycle just as power was lost, the CPU glitched the address lines as it was losing its brains, and the flash corrupted a random location.
Very, very bad.
If a power loss causes the flash to scribble on the wrong SSD location (e.g. the tables that keep track of good and bad blocks), the SSD "dies".
The problem isn't that data is corrupted in the flash, the problem is that the devices' own firmware (the SSD's embedded controller, which is a tiny computer in itself, "boots" from that) is stored in the same memory used for data storage. They could've gotten around this by not storing the firmware on the actual NAND used for data, but a separate device (or kept it inside the controller itself), so any power loss may cause data corruption, but not render the SSD completely unresponsible and inoperable.
That the storage location where a write was in progress is indeterminate afterwards shouldn't matter - between the time that some software initiates a write() and the time an fsync() on the same file returns, there is no guarantee what the written location will contain after a power failure, and if your software relies on the value in any way whatsoever, your software is broken.
The problem is that the flash has an internal state machine that performs the charge tunneling as an iterative process: it tunnels some charge, checks the level on the floating gate, and repeats as necessary. If the power to the flash chip goes away or glitches during this internal programming process, the flash write fails in indeterminate and sometimes very unexpected ways.
So, yes, you of course have to have some low pass in the power supply rail to make sure that power drops no faster than you can handle shutting down the circuit in an orderly fashion - all I am saying is that there is no need to guarantee that a read of a region where a yet-unacknowledged write was happening when the power supply failed returns non-random data, so it is perfectly fine to interrupt the programming process and leave cells where user data is stored in an indeterminate state. It's not OK to glitch address lines while programming is still going on, of course :-)
Most reasonable SSDs do not write cache at all, but thanks to the wear-leveling issues, they need to have a sector-mapping table to keep track of where each sector actually lives. That table takes many, many updates, and since it's usually stored in some form of a tree, it's expensive to save to media which does not support directly overwriting data (IE NAND, which requires a relatively long erase operation to become writable.) This table is typically what is lost during power events, and it is not usually written out when you sync a write.
So what happens is you write, sync, get an ack, lose power, reboot, and magically that sync'd data is either corrupt, or, even worse, it's regained the value it had before your last write, with no indication that there is a problem. This can cause some extremely interesting bugs.
Hu? What kind of applications do you write where 10 ms is too much latency?!
10ms per write is huge. Keep in mind that even a tiny 1KB database insert/update will involve multiple writes: the filesystem journal, the database journal, and the data itself. 10ms for each of those steps adds up quickly.Alternately, consider writing a modestly-sized file to disk, like a 5MB mp3 file. You're spanning multiple flash cells (probably 40 128kb cells at a bare minimum, plus filesystem journaling etc) at that point. Now you're close to half a second of total latency. Oh, you're writing an album's worth of .mp3 files? We're up to five seconds of latency now. But probably more like ten seconds, if your computer is doing anything else whatsoever that involves disk writes in the background.
So yeah, 10ms write latency is no fun.
That you need multiple serialized writes for a commit is a valid point, but I would think that for most applications even a write transaction latency of 100 ms isn't a problem. Also, if you overwrite contents of an existing file without changing its size, you don't need any FS journaling at all, as the FS only needs to maintain metadata consistency, it's not a file-content transaction layer (details depend on the FS, obviously).
Cheap SD cards aren't slow because they have a 10 ms write latency, but because they have a very low IOP rate, which you also could not change by adding a cache and a buffer capacitor, but only by parallelizing writes to flash cells.
There are some recovery strategies employed by the SSD firmware but they can't handle all possible scenarios and you are likely to lose some data in case of an unprotected power failure.
Also, as mentioned before the write takes time and not finishing it on time will make the location being written to have unreadable data.
Corrupting data or BMTs if powered off in the middle of a write is not such a bad thing compared to if that data happens to be the drive's firmware. In the former case at worst you lose a superblock and the OS doesn't boot, but you can still recover from that fairly easily compared to the latter case, where the drive can become no longer responsive.
This would result in a drive with fantastic throughput, except when immediately reading what was written. So long as there was a delay between writes and reads of the same data, it would look like a perfectly performing SSD. This could have fantastic properties for web applications. Given network latencies, many web applications could have considerable delays between the time data is written to disk and subsequently read.
Also, I believe you could also achieve this at a software level with ZFS, which I understand allows you to do things like allocate an entire SSD (or multiple SSDs) as caches for other (presumably slower) drives.
With journalling?
Also, I believe you could also achieve this at a software level with ZFS, which I understand allows you to do things like allocate an entire SSD (or multiple SSDs) as caches for other (presumably slower) drives.
Careful reading, please. Journalling combined with the high throughput of spinning platters if you can create situations that eliminate seek and rotational latency is the key point here, not caching.
You want to stick a spinning rust platter with moving parts in front of a solid-state memory that is several hundred percent faster for sequential tasks, and is an order of magnitude faster for random tasks?
Okay.
A few years ago my employer worked with Intel on a joint IC project (not Flash). My overall impression was that Intel engineers were meticulous and smart. This internal culture probably applies to many different Intel divisions. So I'm not surprised that Intel SSDs are reliable.
My system was configured as a mirror RAID with two 1TB HDDs. The RAID was created with Intel Rapid Storage technology (my motherboard is Asus p8z77v-pro). The failed SSD was used as a RAID cache configured with Intel Smart Response technology.
I was wary to use SSD directly as a main system drive, because I heard of "BAD_CTX 13F" error which happened with Intel 320 drives. My hope was that if SSD is used as just a cache, then in the case of SSD failure the data still be safe. Since this error reportedly occurs during power outage only, I set up an UPS. But all these precautions have not helped.
Yesterday I surfed web with Google Chrome, and my computer suddenly become unresponsive. At first only Chrome was unusually slow, and other open programs work normally, but in a few minutes the computer was totally freeze. I was forced to press reset, and upon restart Windows automatically entered into non-interruptible "recovery mode". After more then 24 hours the OS reported that "further recovery is impossible" and the RAID become unbootable. The SSD serial number was changed to "BAD_CTX 0000013F", the sign of famous "8mb bug". It is interesting that in my case this bug was not caused by any power outage except when I pressed "reset" button, but I don't think this is count as a power loss.
I take an HDD out of the RAID to connect it to other computer and save critical data, but without any success. At first sight all file system looks correct, and I even manage to copy all recent data files, but when I looked into those files it was total mess. Each file consist of some arbitrary chunks of unrelated files, mixed in random order - a bit of some executable file, several lines of my project source code, followed by chunk of some unknown xml configuration file, followed by random bytes, etc. A total mess.
I still don't understand the reason of such spectacular data corruption. I have three hypotheses: 1. SSD cache sent incorrect data to RAID on write (two month ago I switched SSD cache from "enhanced" to "maximized" mode, in which writes initially goes to SSD and only then to the RAID disks). 2. Intel RAID controller goes crazy due to a program error. 3. Windows corrupt data during non-interruptible "recovery" phase.
The moral is, even Intel SSD with UPS is not safe, and mirror RAID cannot not protect data from such errors.
However, uninterruptible power supplies are usually a better investment than power-loss resilient storage media. The problem is, even if your SSD or hard drive behaves perfectly during a power-loss scenario, your server software may not. Almost every database, filesystem, etc. includes some amount of buffering in memory, because sending every write directly to disk is a performance killer.
Also, the best-case scenario with power-loss resilient media is that your system shuts down cleanly. With a UPS, you can keep the system up until diesel generators kick in, a much better endgame for everybody.
I once asked someone who had worked in the hard drive business what a hard drive would do when power was lost. "Try to park the drive head immediately before it crashes on to the platter," was the immediate response. Trying to flush the cache contents wasn't even remotely on his mind. In practice, losing power while writing to a hard disk does often corrupt sectors-- even sectors that weren't being written to during the power loss incident.
It's good to see that (some) SSDs are at least trying to flush the cache, but you really have to ask yourself: can you really trust the manufacturer's claims? And if you can trust them, can you trust your specific software configuration under this unusual scenario? I think it's just too long a frontier to guard with too few sheriffs. Dude, you're getting a UPS.
And your impression of cheap SSDs is dead, flat wrong. They're cheap - every unnecessary part is left off to save money. And we've all (all of us who pay attention) known for years that SSDs (even some with power fail protection) will lose data (even bits which it has reported to have sync'd) on power loss.
A UPS is not enough, if you need to have your data, you need multiple layers of backup, and an SSD must have some method of writing out voltatile data (mostly internal metadata, not cache) before it shuts down.
Source?
And your impression of cheap SSDs is dead, flat wrong. They're cheap - every unnecessary part is left off to save money. And we've all (all of us who pay attention) known for years that SSDs (even some with power fail protection) will lose data (even bits which it has reported to have sync'd) on power loss.
I think you misread what I wrote. I wrote that I would expect cheap SSDs to "survive a power loss event without being bricked." I did not write that they would retain all data, which seems to be what you are arguing against.
I have heard rumors that some cheap SSDs do not honor the SATA SYNC command. Unfortunately I do not have a reliable source for this theory, do you?
A UPS is not enough, if you need to have your data, you need multiple layers of backup, and an SSD must have some method of writing out voltatile data (mostly internal metadata, not cache) before it shuts down.
I don't think anyone is arguing that a UPS is a replacement for backups.
I.e. the savings on a cheaper drive are not worth the risks.
Backup, always, of course.
I had one that went bad in less than 24 hours (my only SSD that needed a warranty claim). I tried to return it to them - filled out form, still had to call them and talk. In the end the guy suggested it would be better for me (much faster) to return it to Amazon than work the return thru intel.
1. Laptops have a built in UPS incase they're unplugged 2. Servers should have UPS incase they're unplugged or small outtages 3. Desktop drives shouldnt be trusted as the only copy. Though I imagine the data corruption would propagate to backups?
On my notebook I've lost a partition with Crucial M4. OS hanged, I did a hard reset and after reboot discovered data loss.
I can imagine a software-fault causing drive-level problems if the drive has a large cache and a broken fsync, or if the bios does some kind of unsafe hard drive reset very quickly after starting.
In any case, it's probably more likely to be file-system level reliability you'd need in the face of driver instability.
OCZ drives had that problem. That's not a surprise.
There's a reason why Intel and Samsung drives consistently average 4.5 stars on Newegg, while OCZ drives consistently average 3-stars.
Some (all?) of the 320 drives (pre-2012) had a bug that basically bricked the drive after power loss. See more on google at '8mb bug' and this Intel thread [1]. The existence of this bug, the well-known reliability reputation of Intel, and the sheer size of this sampling number (N=500?) make the distinction in time important. Were all these drives more recent, or is the Intel failure rate, even with buggy firmware, still ~.5% ?
>However, given that deployment of over 500 Intel 320 SSDs has been carried out and only 3 failures observed over several years, it would be reasonable to conclude that Intel S3500s could be trusted long-term as well
Even with "risky" OS and controller combinations, there was an element of probability involved, so most (probably almost all) power loss events would not hit the bug.
Plus the SSD320 was difficult to obtain back then, and reasonable operators upgraded to the firmware version with this bug fixed, so only a small percentage of the units were ever even vulnerable.
20GB = Twenty Gigabytes
NOT
20gb = Twenty gram bits
and:
20MB/s = Twenty Megabytes per second
NOT
20mbytes/sec = Twenty milli-bytes per second ( If you got you B's and b's correct about you would need to write out "bytes" )
So I refuse to be a wuss and use MiB and GiB.
(Also, get off my lawn.)
In the kilobyte world, 2.4% may not have been too big of a deal.
In the terabyte world, there's a 10% difference between binary and decimal prefixes. That's way bigger than rounding error. We need to start using the binary prefixes.
Kibbles, mibbles, gibbles and tibbles just make me shake my head.
I can actually see where this unit would be useful; in determining the probability of bit rot. The more kelvins you have, and the more bytes you have, the more likely you are to flip a bit due to random fluctuations. temperature*storage capacity = Kelvin-Bytes.
Clearly there will be some variability - and he may have gotten a good, or bad drive - and a larger population of SSDs might behave quite differently in terms of reliability.
I currently have an Samsung 840 1TB and has been rock solid, it replaced a 250GB intel 320. Afaik Samsung is the only manufacturer that owns the whole supply chain, flash + controller.