HPE Drive fail at 32,768 hours without firmware update
support.hpe.com
support.hpe.com
Do they still not test these things with artificially incremented counters?
https://www.engadget.com/2015/05/01/boeing-787-dreamliner-so...
Not that throwing an exception on integer overflow is any better, unless you catch the exception. The classic example here is the Ariane 5 failure:
http://sunnyday.mit.edu/accidents/Ariane5accidentreport.html
"On September 20, 2013, NASA abandoned further attempts to contact the craft.[77] According to chief scientist A'Hearn,[78] the reason for the software malfunction was a Y2K-like problem. August 11, 2013, 00:38:49, was 232 tenth-seconds from January 1, 2000, leading to speculation that a system on the craft tracked time in one-tenth second increments since January 1, 2000, and stored it in an unsigned 32-bit integer, which then overflowed at this time, similar to the Year 2038 problem"
https://en.wikipedia.org/wiki/Deep_Impact_(spacecraft)#Conta...
>Most importantly, the company's already working on an update that will patch the software vulnerability -- though there's no word on when its jets will receive it.
My search of DDG turned up nothing about a resolution. Anyone know?
I know what I would recommend, but marketing would not like it ;-)
Finally, I found an obscure forum post telling me about a firmware bug happening at ~5K hours of disk usage. I updated the firmware and haven't had an issue since.
That's just 71 days of uptime and they hang. There are tens of thousands of these drives deployed as well.
It was a pretty maddening thing to debug and figure out where the issue was (servers, rack, drives, RAID do controllers). 2 different machines 2 and 3 drives. Week later we found the Intel bulletin about the issue.
Thank God for pgbackrest backups.
It was certainly very inconvenient having to reboot systems while we waited for a fix to exist, fortunately we didn't lose too many disks before the fault was identified, which reduced the man hours involved in the DCs.
I'm not sure if any part of this is a scam. A bug, certainly.
http://www.stbsuite.com/support/virtual-training-center/powe...
If you look elsewhere on the Internet, you'll find people with very old and working HDDs that have rolled over, so I suspect this bug is limited to a small number of drives.
(What that page says about not being able to reset it is... not true.)
Likewise, I'm skeptical of "neither the SSD nor the data can be recovered" --- they just want you to buy a new one.
Tangentially related, I wonder how many modern cars will stop working once the odometer rolls over.
If the firmware crashes during boot with negative hour counter, it probably could be only fixed by manually flashing new firmware over JTAG.
Since I run a SMART test every month it is easy to track the hourly progression (and thus rollover) in the event log as the events are reported in POH timing.
Regarding recovery: The FTL is likely toast, in which case while the data probably is unharmed and there, it's basically a giant block-sized jigsaw puzzle. With enough effort, and all the stars align - sure, you might be able to recover some/all of it.
Regarding un-bricking/reset: Potentially, no longer any access to wear-levels at the time. So the future integrity/reliability is kind-of dubious.
If you have, say, a 10-drive wide RAID6 you would need to source drives from 5 manufacturers/batches/models in order to be resilient to that kind of failure. Even if that was feasible that seems horrible to maintain long-term.
Doing a red/blue setup where your red systems use one type of drive and your blue systems use another type of drive seems like it could be reasonably accomplished.
If anything, it's easier to maintain, as all you need to ensure on replacing a drive is to not unintentionally make the array have too many of one type of drive. In practice, it means you just regularly cycle what model you buy for your spares instead of the often totally counter-productive practice of making extra effort to find a supply of the exact same model.
In effect, most places I've done this, it has simply translated into refilling our spares from the currently most cost-effective model or two, and cycling manufacturers, instead of continuing to buy the same model.
The point is not to religiously prevent any kind of potentially unfortunate mixing, because these errors are fairly rare, but to reduce a very real chance using very simple means.
Over the 20+ years I've been doing this, I've seen at least 4-5 cases where homogenous raid arrays have been a major liability (the first one, that taught me to avoid this was the infamous IBM Death Star, where the film on the platters was almost totally scraped off; we had an array that we thankfully didn't lose data from, thanks to backups and careful management once the drives started failing once a week - only for it to take the array 4-5 days to rebuild... we didn't lose data, but we lost a lot of time babysitting that system and working around our dependency on it as a precaution).
I started mixing manufacturers after having had near-misses with several arrays with OCZ drives, where it appears to have been firmware problems across drive models.
> Doing a red/blue setup where your red systems use one type of drive and your blue systems use another type of drive seems like it could be reasonably accomplished.
You need multiple systems too, but the point is that every hour a system is down because of an easily avoidable problem is an hour where your system has reduced resilience and capacity. It's trivial to prevent these kinds of errors from taking down a raid array, so it's pretty pointless not to.
Note that it wouldnțt help in this instance, as the bug is caused by the amount of time a drive was running. Different manufacturers would work, yes.
Which were probably installed and started up at nearly the same time. Oops.
This bug has the potential of simultaneously damaging whole sets of servers, if they were bought and installed in bulk. Dark day indeed.
Have an “off on weekends” node. And 24/7 nodes.
This seems incredibly rich. If you have a bunch of this kit, and you don't immediately shut it down to apply firmware updates, then HPE wash their hands of the consequences.
If you use disks from the same batch in a RAID, they would all begin to fail around the same time, because all of them have the same lifetime more or less.
That's a curious bit of context. It seems to imply they're shifting some of the blame onto their manufacturer? I makes me wonder if this firmware is 100% HPE specific, or if there a 2^16 hours bug about to bite a bunch of other pipelines.
Some advice I've read is to only use unsigned integers if you want to explicitly opt-in to having overflow be defined behavior.
I read somewhere that the true reason for that advice is that it allows the compiler to silently store an "int" loop counter in a 64-bit register, without having to care about 32-bit overflow. If you use size_t for the loop counter, that's no longer an issue.
It commonly factors in loop analysis around unrolling and vectorization, e.g. the loop will run exactly 16*n times OR an integer will overflow and it'll run some other number of times.
Use of signed precludes the overflow and the exact bound enables efficient vectorization.
The difference is that a program which invokes 'implementation defined' behavior can be well defined, whereas a program that invokes undefined behavior is literally free to do anything.
And also more useful warnings from static analysis, since if the analysis can prove that the value will overflow this is guaranteed to be an error.
Our industry really sucks. We need languages where this can't happen, and we need testing procedures where these things are caught. I wonder if software is the industry with the lowest quality:importance ratio.
3,4,and 5-Year 24x7 Carepaqs are available on purchase for the _whole_ chassis and parts... but I wonder if I can keep purchasing 1-year carepaq extensions beyond the 6th year.
I didn't choose YOUR sub-vendor. You did. It's your responsibility to ensure that sub-vendor is operating at your standards. Passing blame to a sub-vendor indicates an unwillingness to take accountability.
I mean, that's no shock coming from HP.
Looks like some sort of run time stored in a signed 2 byte integer. Oops.
Yes, this means that a field meant for diagnosing failures was responsible for a failure. Oops.
If only SSD vendors would do the usual cost-cutting measure of loading firmware from the host computer, this could be trivially fixed.
https://fwupd.org/lvfs/vendors/ https://fwupd.org/lvfs/devices/
The system is quite a lot older than fwupd and less flakey usually. Google for hpsum or HP SPP
It looks like such a bug isn't necessarily SSD specific if it completely bricks the drive.
And while 32.768 hours may seem like a long time for a drive, it's under 4 years of continuous operation. Not unheard of if used in a NAS.
Maybe. OTOH, plenty of people have been running spinning rust drives with way more than 4 years of power-on operation - if this bricking bug was common there, I'm pretty sure we would've noticed. SSD's are a newer tech and it's more common to replace them anyway as specs improve.
SSDs have insanely complex firmware when compared to HDDs so letting this kind of bug slip through is an easier mistake to make in their case.
It doesn't need to be continuous. Total operation is the metric.
/s
I remember when I saw my first ATX computer, and it turned itself off. That was cool.
Our Sun workstations were very stable though.
How is this work legally? For one, how would HPE prove that the customer read the bulletin? I don't imagine they're sending these out via certified mail.
Maybe HP and HPE were more tightly connected 3 years, 270 days 8 hours ago.