Why I'm usually unnerved when modern SSDs die on us
utcc.utoronto.ca
utcc.utoronto.ca
Early flash used to fairly reliable with almost minimal error correction. However with increasing density, smaller processes and multi level cells, it has gone progressively less reliable and slower. Here are some of the things that we need to worry about: https://www.flashmemorysummit.com/English/Collaterals/Procee...
To compensate for all these deficiencies, the SSD architecture and hence the entire FTL becomes very complicated because any part of it can become damaged at any time. We always have to have backup algorithms to recovery from any scenario. Its difficult to build algorithms that can recovery from arbitrary failures in a reasonable time. I cannot have a drive sitting around for 20 minutes trying to fsck itself.
Another problem is that the job while rewarding is not very lucrative. The chance of a multi million dollar payoff for an employee is low. I have a higher chance working on a web connected gadget to become a millionaire. So that means it is really hard to recruit those who are top notch programmers who known how to figure out the algorithms, write the code, debug the hardware. Most new grads these days are interested in python, javascript and machine learning.
> Another problem is that the job while rewarding is not very lucrative.
Do you mean it's lower paying than typical bigcorp software jobs outside of FAANG, or just that there aren't a lot of startups with astronomical valuations in the media FTL space?
The first problem is really two parts, not only is it rare to find people who have passion for storage related technologies but very few will gain exposure to these technologies to develop that passion.
Kids don't routinely grow up with a SAN in the house. They do tend to grow up with lots of internet connected consumer caliber devices and can easily gain exposure to working with these technologies.
I was fortunately able to explore this type of technology in depth because a family owned business let me tinker with their server equipment in high school.
After college I then co-founded a startup back before the cloud became big. That meant we needed to make use of old hardware to provide service to our customers at a price point that made our service profitable. Old drives were not a reliable way to do that. New drives were extremely expensive for old servers back in the day when SCSI was the interface that you expected for a server. We had to get creative and play with JBOD devices. ZFS was an amazing tool for us in those days, and it still is for anyone who wants to tinker.
The other aspect is that while these skills are valuable for creating a "job" they do not have potential for creating "massive wealth". Why learn about storage if you aren't going to be part of the first 10 employees at a company that has a $10B exit? Let Amazon and the other cloud vendors worry about that stuff.
Knowledge is power though. I recently came across an AI startup that I'm now helping. They were spending significant money using GPU computational power to provide artificial intelligence training through a cloud provider. They blew through about $300k in credits within the first year to give you an idea of how much money that type of power can cost.
I am now helping them cut over to their own co-location facility. The first year alone they will save so much money it will pay for the next three years.
Then you read articles like this: https://www.newyorker.com/magazine/2018/12/10/the-friendship...
Reading that helps reinforce the idea that no matter what path you are on in the field there is potential that some random thing you learned about an SSD firmware helps you optimize some growth stage companies product and ultimately that helps you build wealth.
This, exactly.
SSD firmware is opaque, hard to learn from outside. On the other hand, trending web-based framework has all the source code opened, with great documents, and ready to use tools. No wonder young people of today find passion on other things rather than SSD.
That seems like a completely ridiculous way to try to organize your life. Almost no companies have $10B exits.
Just curious how you did that. Did you disassemble their firmware?
Last I checked it was nearly twice as lucrative to be a Ruby-on-rails developer than an embedded engineer.
Embedded also attracts a certain type of engineer, usually very smart and able to manage extreme complexity with attention to detail but at the cost of anything resembling readable, let alone maintainable, software. The fact that anything at all works in the modern world is amazing.
I left the embedded space and have never looked back.
Yeah, yikes. That does not sound appealing :-).
I predict we will lose all of our engineers - embedded dudes included.
I had zero embedded experience when I started doing embedded dev, and then I had zero data experience -- but a surprising amount of general experience is applicable!
Heaven for generalists that always love doing new kinds of things. Once I learned about it, I knew I probably should've done embedded instead of security research. Of course, now there's significant interest in overlap. Might not be too late to learn all that stuff after all. :)
> Heaven for generalists that always love doing new kinds of things.
Really? That's me. I'm in web development, and I've always felt that embedded was far too specialised for me.Though I do prefer readable, maintainable code.
If the drive part of a RAID setup I would actually prefer it just reports itself failed and doesn't slow down access to the array by scanning itself for 20 minutes.
As far as I know, that is one of the main differences when buying enterprise or nas drives compared to consumer drives. With nas drives, the firmware gives up very quickly since the drive is assumed to be part of an array with redundancy. Consumer drives will retry reads for a very long time before reporting i/o error.
When designing a NAND memory product, you aim for some max allowed error rate. You choose the error correction algorithm based on that target. Error rates for NAND products are precisely what the designer intended.
Because SSDs are so large and there is such a large number of them, errors can still occur (at a known rate). FTLs should take that into account. Critical data structures can add checksum redundancy which can reduce the error rate for those to an even lower value, which is usually necessary anyway since power disturbances during erase or program can cause programming errors.
There are of course a number of patterns that increase error rates that FTLs have to be programmed to prevent. Encrypting is the first important step, since it makes the data unlikely to be uniform. The second is mitigating read disturb.
That seems unfortunately true for most low-level infrastructure software. I am very good at those things, and have a medium amount of experience of building web software. Yet I can find a lot more jobs on the latter domain, and they probably won't pay worse (most likely even better).
Besides pay the non-rewarding things in this domain are mostly that the high complexity is often not properly understood by management, rewarded or taken into account in planning. Which then leads to suboptimal products that have been delivered under time pressure.
Manufacturing the hardware requires a lot of capitol to begin with. Why these SSD vendors can't afford the high salary to lure the top talents?
To make it look even worse, the GDP tripled but population only rose 50%, so if you adjusted wages against total economic growth instead of productivity median should be around $110,000. Median.
So yes, $90k is pretty much peanuts for how much money an embedded engineer would be on average making for their employer.
If a line of code is written to run in a customer's browser, then that line of code may be deployed to millions, maybe billions, of customers. But, if an equivalent line of code goes in to firmware for some widget, then you're doing pretty good to get that line in to a million widgets at all, and it's going to take a lot longer too.
You'll still get paid peanuts for doing it.
I always felt that ISPs suffer a similar problem. Nobody cares if everything works as expected, but we'll get upset of it doesn't. However, there is hardly anything outside of what we take for granted they can do that we will actively appreciate. What could a embedded engineer working on SSDs do that will be noticed, appreciated and not taking for granted by customers?
If a software product takes off, that can happen incredibly quickly, and the new product is primarily composed of code. If a hardware product takes off, the change can't be nearly as fast as it's bound by manufacturing, and the code is just one component in each thing.
I happen to be working on firmware for a VOIP phone today - we'll end up making N million of these things, over some number of years. If I were working on an Android app with similar functionality, that app could conceivably go to N million people tomorrow, or 10N, or 100N...
Anyway, I don't think I have a particularly clear or concise (or even correct) argument here, but it's the only way I've been able to rationalise what we've observed.
Ultimately, whatever we write that goes into firmware is hidden from the customer. The customer pays a price per unit, and other hardware vendors are competing against your product. This competition keeps the overall cost low. Except at the very top level like Intel or Samsung, semiconductor manufacturers seem to be fighting one financial crisis after another.
Competition does not work like that in Software. The products (even in the same domain) are all too different from each other, so though they may be competitors, they are rarely in direct competition.
That simple.
"Why do we need source control? It's all there, right on my laptop. Source Depot is just a bunch of trouble." [rough quote from memory, maybe conflated from a couple of engineers]. I'm happy to report that things got better.
On the software side of Xbox the people were much better compensated, and we wrote lots of firmware, too. It was probably harder to be hired, though.
I've threatened to withhold paychecks from employees who have said this to me. The job isn't done till the code is checked in, building, and backed up.
The solution for this would be to open source such projects so engineers from many smaller companies can collaborate. These companies needs to understand that they will not be able to attract top talent and collaboration instead of competition is the way forward.
Regarding your point regarding languages: I would be interested and motivated, but I probably don't have the skill (probably, only did some x86 ASM/C++) nor the location (Europe). Usually someone doesn't just start with C++ but with a managed language, and once someone lands a job becomes demotivated or just doesn't have enough time.
The point being, the solution to this sort of systemic problem is for more corporations to become worker owned. If the guy writing the SSD algorithms has a say in governance and a cut in the profits they will want to stick around. It's the stable version of startups.
The reason you're not getting enough new grads is probably because your compensation isn't competitive enough.
is there any feasible way to recover data after a TRIM command has been issued that you can think of? Is there any way to trick the firmware into not returning 0's when reading the blocks of a deleted file? Mostly interested in doing so for Apple
TRIM destroying the entire data recovery and forensics market seems like such a big deal, I still can't believe it although it started years ago
Why is this a problem?
Let's not forget, we live in a capitalistic society. The job of the capitalists is to exploit labor. The cost of labor is a direct result of the market demand (or not) for a more complex or, by your point, a more robust product.
It need not simply be a market reaction, though. A vendor can create the market for a more complex/robust product. But still, their job is to exploit the available labor so if the people capable of such are available at a lower cost, so be it.
> I have a higher chance working on a web connected gadget to become a millionaire.
As others have noted, perhaps true but the numbers are so small it may as well also be zero. However, as you should now be pointedly aware, the perception doesn't match reality.
Question: do new grads (or let's say up to 3 years experience) working on SSD firmware, making $90k, have to clock in 60+ hours / week and live in insanely high cost regions where their $120k startup salary qualifies them for housing assistance programs and they have to live with at least one if not more roommates? Or is $90k (per another comment) quite good in relation to the total picture?
IOW, my guess is that you are not competing for talent on salary or total comp. My guess is you are competing on general industry attractiveness. The entire mindset around embedded vs web/consumer programming is different. So I think to frame it as a compensation problem is wrong, and by misframing it you will never "solve" it.
Spinning platter drives have parts that form a more relatable metaphor to humans' notions of wear and tear: skates of magnetic readers flying on a cushion of air above a rapidly rotating disc, with the gap separating a few dozen nanometers, often smaller than the process size in the controller's silicon. They have arms that can move the head over a particular disc radius, and a motor that spins the entire stack of platters. These mechanical components exhibit wear proportional to their use -- this makes intuitive sense, and is also recorded in the SMART attributes, so drives in old age and of many park cycles can be replaced preemptively before they catastrophically fail.
SSDs are missing many of the usual mechanisms that would contribute to physical wear leading to sudden catastrophic failure in advanced age. This means that irrespective of their failure rate vs. HDDs, a higher proportion of their catastrophic failures are the fault of the controller. This is discouraging: essentially, the "storage layer" is now quite reliable, so the fallibility of the human-programmed controller is brought into light.
Complete aside, the fly-height of a magnetic head is actually fractions of a nanometer (i.e. hundreds of picometers).
EDIT: I got this from a talk by Bryan Cantrill[1]. The fly-height is allegedly 0.8 nanometers (800 picometers).
Could you provide a reference to a scientific article stating this?
Can you provide some references?
EDIT: There is a paper from 2016 which did an analysis in the difference in the flying height of the head during different operations, and it was measured in Angstroms[2]. I couldn't find one that actually gives a precise value of the flying height. There is a 2011 paper which states that some system they were testing allowed for 4-9 nm flying heights[3] which is about half-an-order-of-magnitude larger than the claim -- but that's already 7 years old.
[1]: https://youtu.be/fE2KDzZaxvE?t=1551 [2]: http://iieng.org/images/proceedings_pdf/E0116009.pdf [3]: http://maeresearch.ucsd.edu/callafon/publications/2011/UweIE...
The read head is something like a jumbo jet flying a handful of feet above the (perfectly smooth) ground. It's really crazy how close these things are, moving very fast. And why accelerometers are a significant feature.[0]
> Can you provide some references?
Wikipedia claims[1]:
> In 2011, the flying height in modern drives was a few nanometers.
and
> The "flying height" is constantly decreasing to enable higher areal density.
> At 7,200 RPM, the edge of the platter is traveling at over 120 kilometres per hour (75 mph)
So you've got a read head flying at 120 km/h => 33 m/s => 33,000,000,000 nm/s at a height of 3nm or less. Picture that!
E.g, a 757 typically cruises at 858 km/h => 238 m/s. So picture your 757 flying at an altitude of 21 nm and that's the metaphor, kinda. The read head is a bit smaller than 1/7 of a 757 jet, obviously.
I am dubious because this is a pretty incredible feat. This is the length scale at which atom-atom interactions become important. That implies that the crystal lattice structure of both the read/write head and the underlying platter will affect the dynamics of the system!
They have to be pretty dang close at 7200+ rpm.
1. I can't find a source that says less than a few nanometers.
2. 300 picometers is roughly the diameter of a helium diatom. The head cannot possibly float through hydrodynamic means if an air molecule can barely even fit under it.
It can. Since siblings liked airplane analogies, here is another one: Consider the head to be an airplane. It has somewhat wing-similar features which provide a lifting force, but the actual read/write head sits below those features (like, say, a landing gear is below wings).
https://en.wikipedia.org/wiki/Atomic_radii_of_the_elements_(...
What we have is that the software is currently less reliable than the memory. There is no fundamental reason for that, it's just that manufacturers put a huge amount of engineering work on reducing the wear, and not so much on programming practices.
These days most of what we call a computer could be described this way. Even your compiled machine language is ultimately far more abstracted from what the processor actually does than it was on, say, a 6502.
Funny that you bring up 6502; that reminds me of 1541 disk drive for C64, which had mostly same 6502 as the host computer (albeit running at slower speed).
It might be an interesting exercise to see how many peripherals are connected to your PC right now that have much more computing power than an 1Mhz 6502.
Actually the spindle isn't touching anything any more, because the spindle/rotor (one part) is supported by a fluid bearing; it basically floats on a thin layer of oil. If the spindle touches the bearing at, essentially any speed that isn't zero, the bearing surfaces are damaged instantly and the resulting burrs and debris will degrade and lock up the bearing very quickly.
I believe the only rolling-element/contact bearing used in modern disks is the pivot bearing of the arm assembly.
So the whole story of a disk being a computer has been true for a long time.
Why shouldn't it? Isn't it just hardware too?
"With spinning HDs, drives might die abruptly but you could at least construct narratives about what could have happened to do that"
Why can't you do the same with SSDs?
It feels like the author's main complaint is the frustration of not understanding SSD hardware as well.
Is this a valid complaint? Are SSDs magical in some way? I'm not an expert but... It's just hardware with pieces that do stuff. Why can't we come up with an understanding of why it fails?
So the combination of non-moving parts (making it hard/impossible to debug via physical inspection) combined with a tons of wear leveling/misc magic can defo make it seem like SSD's are magical.
https://www.youtube.com/watch?v=C5JoC4-qsO0 https://www.forensicswiki.org/wiki/Solid_State_Drive_(SSD)_F...
ZFS has shown that removing layers of abstraction with regard to storage can be beneficial.
It would be an interesting product to have, a raw block API to a Flash device with all the temporary state stored on the host - but a hard one to sell, as it's not differentiated in any way.
Besides, I don't think manufacturers want to release the best practices for using their memory.
Because they don't die incrementally. Wtih a hard disk you'll get bad sectors, growing slowly over time. Or a head crash and then it's all dead.
What could cause an entire SSD to die at once? I would totally understand bad sectors, but the whole thing at once? Where it doesn't even try to read existing data?
Firmware -- undebuggable, unobservable, unfixable software, jammed into your devices -- is the enemy.
A component failing? Electrical components fail. Sometimes it's a manufacturing defect, sometimes a design defect, sometimes something under or over-volted and it was enough to cause damage to any given component.
Could be an IC, could be a capacitor, could be a poorly laid trace. A poorly shielded RF source could even damage any number of components.
I mean, in theory a single high charge particle from that rare cosmic ray that reaches the surface of earth running full-steam-ahead through an IC could cause just the right amount of damage to make it fail although this would be an incredibly improbable scenario.
Same goes for HDDs, televisions, your clock radio, whatever.
SSD has a problem in that almost every failure case looks more like a head-scratch than a bad sector.
Even the most local problem that is excessive writing on a sector can take the entire chip down.
That wire is the one delivering power. Point being, while things are so complicated that you often can have a lot going on and still try things, there are still single points of failure that can never be fully covered.
Indeed, that is the question that the article's author seems to be asking.
What is so frustrating about SSDs is how very poorly they compare to previous incarnations of solid state storage.
Using Disk-On-Chip and/or IDE-pin-compatible CF cards, I had many, many devices in the field that lasted, mounted read-only, for decades An entire sect of the computing industry came to rely on these parts as alternatives to spinning media that could not mechanically fail.
This is not the case with SSDs at all. They fail left and right, even mounted read-only, for all manner of complicated and interesting reasons. It's very frustrating that SSDs are not a step forward in reliability from spinning media and are a step downward compared to (for instance) a 16MB consumer CF card from Sandisk, circa 2000.
rsync.net filers, which need a boot mirror, are always constructed with two unrelated SSDs - usually one Intel part and one Samsung part - so that when the inevitable usage-related failure occurs, it does not occur simultaneously to both members of the mirror which have, being a mirror, been subjected to identical usage-lives.[1]
We shouldn't have to do that.
[1] I can't overstate this - if you need a RAID mirror, do not use identical SSDs for the two members of the mirror. There are many, many cases of SSDs failing not due to "wear" or end-of-life, but due to weird usage edge cases that cause them to puke ... and in a mirror, you give the two parts identical usage ... we either get two different generations of Intel part (current gen and just-previous gen) or we get current Intel and current Samsung ...
Also, though, because most serious RAIDs contain more drives than you can find manufacturers of hard drives.
The tradeoff from worse read disturb characteristics is NAND that is 100x cheaper per GB than in 2000.
But magnetic hard drives failed all the time. I have a giant stack in my office closet just from my dev machines over the years. But it wasn't new and scary -- it was just a hard drive failing -- so it was just normal. Some had controllers fail, suddenly blinking out of existence. Another had cache memory corrupt so it just gave ridiculous readings occasionally. Others had physical failures.
I don't know where to begin relative to prior flash memory (e.g. CF cards) which were absolutely notorious trash.
It is worth noting that every smartphone the world over has an "SSD" in it. We spend remarkably little of our mental power concerned about the flash storage. It is the cause of a negligible amount of device failures.
That would almost certainly be small-block SLC flash rated for 100K program/erase cycles and 10 years of retention. The huge-block TLC now is ~1K program/erase cycles and 2-3 years of of retention depending on where you look (the manufacturers are not surprisingly quite reluctant to release these specs...)
I don't know about that. I know most managers in embedded software go cheap on engineers and software assurance on purpose to get bonuses and such. The hardware side makes me think you haven't studied deep, sub-micron hardware much. I started looking into it a few years ago or so, just reading the slides on lots of stuff even though not understanding much of it. They helpfully put a lot in lay terms, though, with lots of comparisons. To say the stuff gets harder every time you shrink to a smaller node is an understatement, esp for solid state.
If anything, the modern flash should probably be considered broken right as they ship out of the factory. If not, the process nodes after 90nm or so just keep having more and more ways for individual components to screw up or change behavior across same wafer. Some happen instantly by design. Some happen later with aging. The memory technologies are closer to that analog level of things than most with harder verification. The high-density flash on newer nodes uses less-reliable tech than most just to operate at that low cost. So, they add all kinds of firmware and hardware tricks to try to make it work for a period of time like it's a whole, functional unit of storage despite pieces of it misbehaving all throughout.
It's a nice, man-made miracle these techs even work at all. Those that last longer like you mentioned still exist. I'll add a comment with links to one type so you can compare price/storage/performance to these broken-by-design SSD's you use. I'll throw in another two that mention shrinking challenges so you can see what they face every time they have to upgrade or just deploy new designs in mixed-signal.
https://news.ycombinator.com/item?id=18528023
https://www.electronicdesign.com/digital-ics/understanding-2...
https://anysilicon.com/ic-design-impact-in-moving-from-28nm-...
In a mechanical hard drive, there are moving parts which can wear out due to friction, etc.
SSDs are solid-state, so it seems like at least theoretically, it should be possible to build one that keeps working for decades. e.g. I have solid-state hardware from the 70s and 80s which still functions.
I've always been a little mystified as to why SSDs' data areas wear out for that reason,[1] but that's a whole separate issue. The author of the article is writing about sudden failure of the device as a whole.
There are only two explanations I've ever heard for short lifetimes in electronics as a general industry.[2]
For devices manufactured after about 2000 there's tin whiskers,[3] which began to be a problem because of RoHS requirements. I'm not sure if that applies here, though.
If the device includes electrolytic capacitors, my understanding is that those generally have a finite lifetime as well, and it can be fairly short if they're poorly-made.
I'm not a hardware expert, though, so I'd be interested in hearing about other factors, and I'm sure the author of the article would too.
[1] I've seen lots of writeups of how wear-leveling works, etc., but never a good physical explanation of what is actually wearing out over time.
[2] Obviously there are other factors for specific devices, or specific designs. E.g. maybe parts of a specific device break over time due to thermal expansion and contraction if the device wasn't engineered to handle that properly.
there's a reason companies don't repair circuit boards in consumer electronics.
It doesn't seem plausible to have a business case where you'd get be able to get a large quantity of identical boards (that are otherwise good!), replace caps on them, and be able to sell them for much more than you got them - i.e. that the boards haven't become obsolete in that time. If there's no mass production, there's not much use for automation.
EDIT: More reasons and methods for failure of "solid state" circuits and systems: https://news.ycombinator.com/item?id=14765868
SSDs are flash are basically EEPROMs. In an EEPROM one bit is stored in a dual-gate MOSFET. One Gate is a normal gate, the other is floating, i.e. it just is a small conductive island. The information is stored by (quite literally) shooting electrons through the insulation of the floating gate into it. They're then trapped on the gate; if you turn the second gate on, the transistor conducts iff the floating gate is also turned on. This shooting action happens to damage the insulation, which at some point is degraded enough that it can't keep the electrons trapped on the floating gate. Hence the electrons leave, together with your information.
> For devices manufactured after about 2000 there's tin whiskers,[3] which began to be a problem because of RoHS requirements.
The solder joints themselves usually don't form whiskers, which mostly grow on pure tin-plated surfaces, e.g. the pins of components (which previously used lead). Using lead-free solder is mostly co-incident with whisker risk, not the cause of the majority of problems.
If I keep my SSD 40 degrees cooler with a freezer, will it last an indefinite number of write cycles?
That defect doesn't strike me as being inherently related to SSD media, but really left a bad taste in my mouth with what might be going on in the development process to lead to such instability.
I wonder if there’s an easy way to distinguish a controller failure from a flash failure from the behavior of the device over the last few seconds/minutes of operation. In theory a controller failure should cause a fairly abrupt loss of service, but I’m sure there are soft lockup failure modes too.
With SSD or NVMe the controller isn't really a separate component you can just replace. Maybe it's possible to saw off the broken part and bodge-wire it to a working surrogate, but that would be extreme.
A tear-down of a broken SSD might reveal more about what could be done.
I havent run into this issue for ~4 years, so I'm not sure if this is a solved problem or I got lucky a few years back.
I haven't had a dead SSD problem in a while though, so I don't know how common firmware issues are now.
Sometimes the dead SSDs will respond to a handful of commands anyway, meaning that you can attempt a firmware upgrade or reset. That’s the lucky case. But in the majority of cases I’ve seen, the firmware/controller goes into some hard lockup where it no longer processes SATA commands at all. I wish these things had JTAG ports...
I don't know. I'm just trying to say that since flash isn't volatile the SLC cache isn't treated as volatile cache either.
OCZ does not exist anymore. Not sure of if this was one of the causes, but I would otherwise never buy a drive from them again.
hot-air rework to lift all the flash chips off and get them hooked up to something (probably custom) that can read them.
if the controller was encrypting everything that went to flash, you also get to try and find the key in the controller's memory.
your data is probably jumbled up one way or the other, but at least a custom board will let you read all of the underlying flash, instead of just the portions of it that a new controller would believe are in use. (keep in mind that the SSDs will have more physical blocks than they advertise logically.)
A lot of SSDs also store their main firmware in the same flash that is used to hold user data... this is something which was done with hard drives too (hence why dead/dying HDDs sometimes show up as a small drive with a weird name --- that's the "recovery mode").
Needless to say, there are no mysterious storage gods. These are artifacts made by humans, and somewhere out there, there is an engineer who either understands why these failures are happening, or knows how to engineer these devices in such a way that when these failures happen, the cause can be determined, and then a design iteration can be done to reduce the failure rate and make the failure modes more robust. The reason this doesn't happen is that customers aren't demanding it. If major purchasers started demanding, essentially, an SLA from their SSD manufacturers, with actual financial consequences for violating it, you would be amazed how fast all of these problems would get fixed. But instead we vent our frustrations in blog posts and HN comments :-(
What about Consus, the god who protected grain storage in the ancient Roman religion? [1]. Or Eopsin, the Korean goddess of storage? [2]
If you want your storage to have an SLA, storage service providers exist, and will be happy to give you an SLA if you're willing to pay. But it isn't cheap.
Yes, I know that. That's why I said "essentially".
> SSDs have the physical equivalent of an SLA, a warranty.
There are two orthogonal issues. The first is what happens when a device fails. A warranty addresses that. The second is how does it fail. Does it fail all at once with no warning, no way to perform post-mortem diagnostics, and no way to recover the data? Or does it fail with a gradual degradation of performance and capacity over time, and in a way that, if/when total failure occurs, the cause can be ascertained and the data still recovered somehow?
how does that have any relation to an SLA? An SLA is a promise of a certain amount of uptime, with financial penalty for the provider if not met. A warranty is a promise of a certain product lifespan, with a financial penalty for the provider if not met.
SLAs have nothing to do with providing diagnostic information.
I had a 512GB Samsung drive that became very slow randomly at doing IO operations, the whole machine would die for 10-30 seconds at a time once or twice a day while any process that tried to use the disk became blocked on IO. Then it'd come right back like everything was perfectly fine.
Issues like this definitely worry me, we're basically completely blind as to what those controllers and flash chips are actually doing. Not that it wasn't a similar situation with HDD controllers before, but at least it didn't seem as unpredictable.
In the past 3 months, there's been maybe 5 times where I started my laptop and it took 5+ minutes to finish booting. Normally, it's 15 seconds. Once I'm logged in, doing anything takes forever, but it does eventually load. A power off, and back on got it "working" again but who knows for how long.
I've had this drive for 4-5 years now, so I'm impressed it's lasted this long.
As densities of data get higher and higher, it doesn't take much to have a catastrophic data failure. The only way to protect against this is having multiple replicas of your data.
OS: Yes, I did.
OS' inner voice: Nah, I didn't, but I'll do it soon, I think.
OS to drive: Those write commands, they're durable yet?
Drive: Sure!
Drive's inner voice: Nah, they're still in my RAM, but I'll probably write them out any time now.
Because they've shown most metrics are kinda useless or mean different things from one manufacturer to the other.
https://www.backblaze.com/blog/what-smart-stats-indicate-har... https://www.backblaze.com/blog/hard-drive-smart-stats/
We've experienced exactly the same thing. Our general course of action is to perform a hard power cycle of the server through IPMI - a warm cycle doesn't seem to work. I've always presumed it was down to dodgy SSD controller firmware given the way it suddenly stops appearing in the output of fdisk -l.
It's sort of a devil's bargain - the performance of SSDs is so much better that I can't pass up using it over a spinning disk even if they occasionally lose everything. There was a great game for the original Nintendo called "Pinball Quest". As you advanced through the game you could get upgrades such as side stoppers, stronger flippers, etc. You bought these items from a demon in between levels. After the red "Strong Flippers", the next upgrade was the purple "Devil's Flippers". The trick was that occasionally they'd turn to stone when you needed them and possibly cause you to lose the pinball. But they were such an upgrade over the Strong Flippers (when they weren't turned to stone) that you bought them anyway.
SSDs are kind of like that.
I'm more worried about everybody else who gets an SSD and doesn't take the right precautions, because everybody sells SSDs as being so much more reliable than mechanical drives.
I worked as a PC technician for a while recently. Of the handful of catastrophic failures of mechanical drives we had, the majority of those were ones that were physically dropped, resulting in a head crash. Otherwise, we generally always managed to save data from failing drives. Any failing SSD we encountered was just dead, since there are really only two states: Fine or failed. There was nothing we could do except refer them to a data recovery company that charges thousands of Euros.
You also CAN'T HEAR if there is an issue, which acts as another warning sign that something might be going wrong or will go wrong soon. Loud tickings or clickings or overwork is a sure sign to start backing up and get ready to buy a new drive!
So the citation of some evidence is required to back up your claim.
The author doesn't specify a metric for writes, but based on the MX300 specs that I find, it can go up to 219 GB/day for the 2TB drives he uses, or 87 GB/day for a couple of the 525GB he still had. He doesn't specify though.
I'm a little scared about my new SSDs that have replaced a few rust-spinners in our data center.
Every SSD failure I've had, the failure mode was "what SSD?"
Now, I realise most people should ponder their backup regime before the tictictick, not after. But as the phrase goes "The best time to plant a tree was 20 years ago. The second-best is now." The SSD equivalent is "The be- nope, too late."
They're just terribly unforgiving, which doesn't fit with a culture that values cure over prevention.
Perhaps I'll grow more comfortable after another decade or so, when there is enough real world experience to go by.
I'm of the opinion that harddrives don't actually function, they just maintain the illusion while they wait for a more interesting moment to die.
Who in the modern age doesn't back up everything all the time? Don't we all operate with the assumption these things are going to blow at any time? 90%+ of my data is on cloud storage now anyway. When a SSD goes out don't you just chunk it in the drawer of old drives that you promise to take to the disposal center this weekend (and never do) and then take a quick trip to your local computer store for a new one?
This reminds me of something an IT support staffer told me a long time ago... "The difference between a IT pro and a user is that to an IT pro hard drives are a consumable resources".
Thank goodness for Backblaze, Time Machine, Carbon Copy Cloner, Drobo and Synology. Maybe I have gone overboard, but I have not lost any data in 12+ years.
If we'd be arguing about mean time between failure or total cost of ownership, then it'd be relevant, but this post isn't even claiming that the rates of failure are excessive (compared to what?), just that they are too weird and unpredictable for the author's liking.
Some years ago I got a great deal on several Pacer disks and wrote a program to write a pseudo-random sequence of data (using a known initial seed) across the entire disk and read it back and compare. Part way through, the data didn't match. No ECC errors, nothing raised by the filesystem, just mismatched bits which came back in a manner which tried to "trick" me into thinking they were good data. This happened on like 5 of the 8 disks. Needless to say I sent those crappy SSD's back to the manufacturer (unfortunately only got a 2/3 refund) along with some harsh words for their engineers.
I've had more name-brand SSD's fail, in various manners (even on well-reviewed Kingston drives). Sometimes in ways that can't be accessed at all, other times (at best of times) in a manner which doesn't allow writes but still allows reads (albeit at a trickle of a datarate).
These days I use solely Intel-based, top-line SSD's, and some (very limited) Samsungs. The choice isn't based on empirical data, but rather and impression their bar is a little higher (or more conservative) in terms of reliability, and simply not wanting to deal with the apparent issues I've seemed to encounter with other brands. The downtime lost from restoring / reconstructing just isn't worth it to me. Maybe I'm paying twice as much as I ought to, but since making the switch many years back it's worked out pretty well and I've been happy / fortunate.
I run my SSD's in RAID10 using high-end controllers (aside from a few in ZFS).
Just my own subjective experiences, again I'm not doing this at scale.
(Um, here's where I have to be critical of tarsnap: their recovery performance is absolutely abysmal for small files. They're latency bound between you, their EC2 instance, and the backing S3 store. Think single or double digit kB/s and then think about how much data you back up with tarsnap. I can't recommend any other backup provider better, but this is an experience where tarsnap left me very disappointed.)
Looking at that SSD and my other SSDs' SMART data, they report extra blocks remaining in SMART, and you can monitor that as it goes down. Ideally you replace the drive before it gets to zero.
My primary mistake was simply not monitoring that data in an effective way.
I don't think anyone who monitors HDDs has any real expectation that the high-level SMART yes/no is going to protect them from data loss. Instead they look at highly predictive factors like "Reallocated_Sector_Ct" or "Raw_Read_Error_Rate" (or even plain old "Power_On_Hours").
For SSDs it's quite similar: Reallocated_Sector_Ct, Power_On_Hours_and_Msec, Available_Reservd_Space, Uncorrectable_Error_Cnt, Erase_Fail_Count, Workld_Media_Wear_Indic, Media_Wearout_Indicator. Maybe NAND_Writes_1GiB.
NVME SSDs provide SMART-like data on logpage 2 ("Available spare", "Percentage used", "Power on hours"). For some reason the NVME spec does not require media to accept host-initiated self-checks, so most NVMe drives don't have the same functionality as smartctl --test. :-(
Not too much later, I used Iomega ZIP drives, and experienced the "click of death". That was sudden, and irreversible, but also very understandable.
For the past couple decades, I've consistently used RAID arrays, mostly RAID1 or RAID10 (and RAID0 or RAID5-6 for ephemeral stuff). I've had several HDD failures, but they were usually progressive, and I just swapped out and rebuilt.
I recently had my first SSD failure. And it was also progressive. The first symptom was system freeze, requiring hard reboot, and then I'd see that one of the SSDs had dropped out of the array. But I could add it back. At first, I thought that there was some software problem, and that the RAID stuff was just caused by hard reboot.
But eventually, the box wouldn't boot, so I had to replace the bad SSD and rebuilt the array. It was complicated by having sd1 RAID10 for /boot, and sd5 RAID10 for LVM2 and LUKS. So I also had to run fdisk before device mapper would work.
Just like not all hard drives are created equal. My previous job involved a decade running 10 cabinets of servers an hour away with very little manpower: we eventually came to find that IBM/HGST drives were a lot more reliable than others.
We also evaluated some early SSDs, and they were terribly unreliable. We eventually settled on the Intel drives and they were superb. My new job we've been using mostly Intel and Samsung Pro drives, they work great. But Dell sent us a server with some "enterprise SSDs" in it, that we eventually found were Plextor drives. Those things were terrible. We replaced them immediately with Intel, but used some of the Plextor drives and had all of them fail within a year. I'd put the Intel 64GB SLC drives from our 7 year old database server in a system before I'd put one of those brand new "enterprise" Plextor drives in.
I love Crucial, I buy a lot of RAM from them, but I'm skeptical of switching to other brands of SSDs. The more experience I have, the more conservative I get with systems that matter.
http://www.assetinsights.net/Glossary/G_Elegant_Degradation....
http://www.assetinsights.net/Glossary/G_Graceful_Degradation...
I am loathe to even keep my personal data at home on one drive and since I use an iMac that requires me to have time machine as mirroring/etc of the internal drive is not truly possible; at least I did not spend enough time researching it
With the older drives you sometimes would have a drive die, replace it, restore your backup only to find that in the process of dying the drive was actually corrupting some of the data which went into the backups, now you've got to hunt down the last uncorrupted versions of the data in the backup…
Never tried baking a drive (ssd or hd), but I have with a red ring'd xbox360 mobo.
Sometimes the "low tech" solutions still work.
I've also had some hard drives that you could bring back to life by giving them a firm knock with your knuckles too.
Doesn't really answer much, but it's a last ditch effort that has saved me more times than not.
We have all sorts of knowledge about about it but when something happens we're still looking for an explanation for each instance.
If you think of it like nuclear decay you'll still be able to say things about the ensemble, but not each individual member.
This is incorrect. As much of the argument seems predicated on this I don't see a real issue.
What are some best practices for personal hard drive crash warly-warning?
(I filed a request for a bike in my early life, so never got to that 100).
I use arq w/ backblaze