Power Failure Testing with SSDs
blog.nordeus.com
blog.nordeus.com
Disclosure: we run Postgresql in production on Intel and Samsung "data center grade" SSDs, and I participated in the aforementioned PG mailing list discussions.
Updated: from http://www.postgresql.org/docs/9.4/static/wal-reliability.ht...
"Many solid-state drives (SSD) also have volatile write-back caches."
and this thread: http://www.postgresql.org/message-id/533F405F.2050106@benjam...
Also I know it may sound funny to use these SSDs in production, but when we started with our game Top Eleven we didn't know that, when you have 5 people working on a game, you don't have time (nor resources) to think about a type of an SSD; you order server with an SSD, you get an SSD (you usually don't even have a choice when you rent) and you use it. This is how it still works with almost any server renting company.
What was even more interesting is that our SSDs are connected to the Dell H700/H710 RAID controller which has a battery backup unit (BBU) which should make our drives power failure resilient. RAID controller with BBU in case of a power failure can hold the cached data until the power comes back, so that it can flush it to the drives when the drives come back online.
RAID controllers that you could afford for yourself, costing less than $1000, are just another point of failure. I've seen them fail more often than the disks, causing corruption as they went. However I've also seen the very expensive datacenter RAID controllers keep a whole bunch of servers up for years as their (15k rpm spinning rust) disks were failing and being swapped.
SSD problems with power failure are such that ZFS won't save you. ZFS's checksumming will at least know what's corrupted and what isn't. But the extent of the corruption will be practically unlimited if the OS can't trust the drive to honor write barriers / flushes - any byte since the last powerup could still be only in volatile cache on the SSD and never make it to persistent NAND.
Every filesystem depends on these occasional flushes/barriers to establish checkpoints where all previous writes are really written, including ZFS. Consider, it may have created an updated copy of the filesystem tree root node, made sure it was flushed, made another updated copy, made sure it was flushed, and then (only after the flushes) re-used the space that the original copy occupied for other data. If you can't trust that the flushes actually happened when the drive indicated they were completed, then it might be that only the last write, over the original root node copy, actually made it to the NAND.
As for whether or not it's superior to ZFS, though, that's a tricky question. High-end hardware RAID gets you a lot of neat features that you don't find elsewhere, but it's expensive and the RAID controller itself tends to become more and more of a point of failure (always keep spares, especially if they go out of manufacture!)
The above has absolutely zero to do with the drive write cache and power protection. The drive write cache is used to cache the write on the drive after the RAID controller. Spinning metal drives have caches. SSD's have caches. Whether you use the cache is up to you and your OS. The safest setting is off. You can usually get more performance by leaving it on. This is why Windows would throw warnings all over the place when enabling write caching without a connected UPS. I'm not sure exactly what the logic is on Linux for sd/sg devices.
The default Dell RAID controller behavior with write caching on drives is based on drive interface. SAS = off. SATA = on. If you run non-power safe SATA drives with a Dell RAID controller, you must disable the write cache if you want data durability.
Then again, you could also run your databases on power-safe, non-consumer SSD's. This whole article was basically a giant warning sign saying "run away they need serious help."
No, there are others that cargo-cult believe in ZFS too
-- a hyped single vendor OS, without first-tier-support on Linux, and with its own issues, compared to a industry wide standard, used for 3 decades in the most demanding data-centers protocol and its implementations.
Who's that single vendor? FreeBSD, OpenIndiana, Joyent, Oracle?
compared to a industry wide standard
Good luck exchanging your fried RAID controller for a "comparable" model
If you are running Linux/*BSD/SmartOS, the best RAID controller is a JBOD controller, i.e. one that that exposes all connected disks as-is to the host OS, and than it would be in-kernel soft-RAID implementing the actual logic.
This approach gives you a more predictable system with no vendor-specific idiosyncrasies or extra cache level to worry about (BBU in OP's article); you end up running RAID code that is peer-reviewed, fully integrated into fsync() algorithm, and surprisingly, often gives you better performance too.
"True hardware RAID controllers", on the other hand, are nothing more than an application-specific computer with its own (non-upgradeable and often outdated) CPU, RAM, I/O and hard-to-upgrade proprietary software.
And if you buy into the view described above, than replacing a JBOD controller for RAID use is exactly the same thing as replacing JBOD controller for ZFS use, by definition.
These drives are defective, and they should refund customers for the SSD prices, plus some hefty compensation for the potential data loss they could have or have caused.
If you make a storage system and ack a write/sync, it had better be durably written.
An HDD will not hold data that is written to its write cache either so the SSDs are well within the spec.
The complaint is not about losing data in a volatile cache the complaint is that drives will lose data even after the drive has claimed to have flushed it's write cache after being given a write barrier.
840Pro is a consumer grade drive and has reduced guaranteed max write capacity and larger variance in quality from unit to unit. Instead you should use server type disks instead: http://www.samsung.com/global/business/semiconductor/minisit... We use these in our analytics servers, they have stood up to test of time many times, and without any issues. SSD tech evolves so fast that 3 years since release of model lineup seems like forever, especially in an intensive write prone environment(for example writing slows reading but by how much? DC level drives are way ahead in this and very consistent as per what we found out in our tryouts).
The "disk cache" is a disk hardware option (how it uses its own RAM), if I understood, and the barriers are just an option of the FS behavior (software). I'd expect that the performance penalty to the former is much higher than to the later?
Usually this works like this: an external voltage powers the drive, inside the drive this voltage is split into two circuit paths. One drives a voltage regulator with a large capacitor on the far side of it, the other drives, usually accompanied by a pull down resistor an input of the cpu in the drive. If that input goes below a specified value then the cpu knows the power is about to fail and will trigger the cache write to the persistent media.
Obviously that only holds for the data already present in the memory of the CPU on the drive, the rest of the computer will have to fend for itself.
So just adding a capacitor on the outside will keep the drive powered up for half a second longer but won't initiate a cache dump (assuming that you're not going to end up powering the rest of the circuitry as well, you need to add a diode or something to avoid backfeeding the circuit that charged the cap in the first place).
Power loss can be very challenging to get robust especially if the controller uses clever algorithms at runtime to get better performance because recovering the drive state from sudden power loss is more difficult. That is why I think the Intel 520 controller has so many problems because they are using a SandForce controller which was known to use no external DRAM and compression algorithms which just complicates things.
A large supercap in the 100-200mF range for an SSD is around $1 or $2. In fact you can implement power loss protection with less capacitance with regular tantalum caps like the Intel 320 did (http://www.storagereview.com/intel_ssd_320_review_300gb). But drive manufacturers see the consumer market doesn't care about power loss protection, so they decide to scrap the feature, which saves a buck or two, and saves some PCB space.
It isn't a clean OS shutdown usually, but an orderly transition to a hibernation state of some variety which should include flushing drives.
I think if you're selling laptops, then you should worry about these types of cases. Not to mention cases like having Windows Updates run on battery which means the laptop can't hibernate when after its started these installs at shutdown.
Standby/hibernate is still far from perfect. A fifty cent capacitor shouldn't be a dealbreaker for ssd manufacturers.
I've recently bought a new SSD and was searching for information on power loss protection, and the only vendor documentation on the matters seems to be [1] and [2]. SSD reviews have plenty of performance numbers (that interest me less), and besides sometimes describing what the vendor says about power loss protection they don't perform any actual testing for power loss protection, and end up being fooled by the vendors sometimes [3]. The only actual test I found was [4].
Capacitors are one way of protecting, but for some reason even some of the newer enterprise SSDs sometimes have them, sometimes not, even if older versions had them. Some SSDs claim to have journaling (on SLC NAND) instead, but given how the firmware is closed source there is no way to inspect it for bugs.
[1] "Storage devices require a graceful removal of power to ensure data integrity is preserved. Graceful removal of power includes commands to signal to the storage device that power might be imminently removed" http://www.sandisk.com/Assets/docs/Unexpected_Power_Loss_Pro...
[2] http://www.intel.com/content/dam/www/public/us/en/documents/...
[3] http://www.anandtech.com/show/8528/micron-m600-128gb-256gb-1...
And the market segmentation is the natural way to go. The different needs and different sensitivity to price determine the price.
Generally power loss will not effect these drives, and I couldn't get them to scramble existing data or damage the drive while unplugging them or unplugging the computer they were in during heavy writes.
M600DC is Crucial's full scale DC model that offers superior power loss protection.
Crucial drives are manufactured at the joint Intel/Micron facility (using technology from both companies) that Intel's current lineup of drives are manufactured.
I agree with the article that S3500s have sufficient protection.
SanDisk also has a power loss protected drive, but the drives themselves don't seem to be any good. I'm hoping SanDisk drives produced under Western Digital's ownership will be much better.
write barriers is a mount option [1] [2].