SSD will fail at 40k power-on hours (2021)
cisco.com
cisco.com
(TLDR For anyone wondering, "recent HN issues" means HN very likely went down yesterday because of this same bug, when two (edit: two pairs, four total) enterprise SSDs with old firmware died after 40,000 hours close together. An admin of HN and its host both like this theory. See details in that thread.)
Edit: If you want to discuss that theory, it's probably better to do it in that other thread directly instead... dang and a person from M5 Hosting (HN's previous host) are both participating there.
It's rare, but it's not _that_ rare. You have to make the effort to understand why it failed.
I had a situation where I deployed Toshiba SLC SSDs (that were purchased over the course of several months) and a piece of software that synchronized to disk frequently, resulting in about 1GB of writes per hour.
After ~11 months in service, most of the drives died in the same 4 week period. We were astounded that everything failed so close to each other, including instances where both drives in a RAID 1 set were toast.
We did extensive troubleshooting between the failed servers and the remaining servers and figured out that write volume (by proxy of in-service date) was the one predictor of failure. Shortly thereafter, wear leveling and TRIM became things we sought out mentions of when spec'ing out hardware.
IBM Deskstar: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
IBM Deathstar: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
May 26, 2006: https://www.pcworld.com/article/535838/worst_products_ever.h...
October 9, 2006: https://news.ycombinator.com/item?id=1
Memory would still have been reasonably green.
Wondering if it's real... https://m.slashdot.org/story/20680
Years later, it's a widespread phenomenon: https://m.slashdot.org/story/43312
It was affectionately called the "click of death".
Ah, inventing an opinion to get angry about: some things never change.
So if you only formatted them (filesystem wise) out to capacity-2Gb they were a really cheap option at the time.
I would not have let a normal business user near one, but the developers I was supporting were most pleased about their larger than expected scratch disks for test databases and intermediate compilation artifacts.
Everything breaks. Things that at least break predictably make me happier than the alternative.
Absolute best mechanical drives available until quite recently can be traced back to Deatstar. Deskstar 7K4000 were absolute best in class.
https://en.wikipedia.org/wiki/Deskstar
Hitachi bought IBM hard drive business in 2003 for $2B. Sadly Its now owned by WD.
HGST SAS SSDs (the ones that pair Intel NAND and Hitachi SAS controllers) have also been reliable performers in my home office experience, even without (RAID controller support for) UNMAP (SCSI "TRIM"). Incidentally, these now appear to be selling on eBay for more than I paid for them "lightly used" several years ago.
Both primary and failover servers had RAID arrays. I suspect RAID 10 (striped mirror), which would mean two drives would have to fail to take down a single server.
Four drives of the same manufacturer spec and batch would do that.
https://news.ycombinator.com/item?id=32031655
I've increasingly come to view systems operations / SRE as a risk management exercise, where the goal is to reduce the odds of a catastrophic failure. Total system outage is one level, unrecoverable total system outage is even worse.
Having multiple redundant backups / storage systems, in different locations, with different vendor hardware / stacks, all helps reduce risk of a single-factor outage. Though complexity risk is its own issue.
s/reduce/find an appropriate level for/
It's a common misconception that risk management and risk reduction are synonyms. Risk management is about finding the right level of risk given external factors. Sometimes that means maintaining the current level of risk or even increasing it in favour of other properties.
Quibbling over whether the proper term is risk management or risk reduction rather spectacularly misses the forest for the trees.
I don't know how long you've been in the business, but the change seems a relatively recent one, one that wasn't manifestly obvious to me, and one that has pretty much always seemed difficult to communicate to management.
Whether that's because business management is often about ignoring risks or treating it as inconvenient, or if I've just had a long string of bad bosses, I'm not sure.
I did make a point of looking through several of the books that were formative for me (mostly 1990s and 2000s publication dates), and there's little addressing the point. Limonchelli's book on time management for sysadmins was a notable departure from the standard when it came out, in 2008. I'd say that marked the shift toward structured and process-oriented practices.
That was about the time of the transition from "pets" to "cattle" (focus on individual servers vs. groups / farms), but pre-dates the cloud transition.
The stuff I've read that touches on this idea is almost all from 2006 and onwards, mainly 2010s. The earliest example is a bit of an outlier: Douglas Hubbard's 1985 How to Measure Anything -- but it's also only tangentially related.
The other real exceptions are books on statistics (where the idea of risk management -- at least in my collection -- seems to have gotten popular in the 1950s, probably as a result of the second World war) and financial risk management (which seems to really have taken off in the 1980s, probably in conjunction with options becoming a thing.) Statisticians and finance people (and by extension e.g. poker and bridge players) have known this stuff for a while.
Of course, hydrologists have been doing this stuff since the early 1900s at least, but extreme value theory has always been a kind of niche so I'm not sure I should count that.
----
That said, I did mention it was obvious to me. I still find it hard to convince management and colleagues of its importance...
I didn’t keep good track of such things but a lot of my early reading in the 70’s was in operations research and decision support systems, mostly sort of what we call operational analytics these days with a big helping of statistical process control too. World War 2 logistics practices and ‘50s and ‘60s “scientific management” fads generated a lot of material, some insightful. Many medium-sized businesses could afford significant R&D then, so you’ll find e.g. furniture factories developing their own computer systems from PCBs to custom ASIC components, just to manage statistical process control and decision support systems.
> That said, I did mention it was obvious to me. I still find it hard to convince management and colleagues of its importance...
I think the reason I keep having to justify this every few years is the tendency towards abstractions in management which try to simplify things into “anecdotal analytics”, e.g. preferring a persuasive narrative over reality...for a good cynical perspective from the ‘50s I recommend C.M. Kornbluth’s “The Marching Morons” (<https://en.wikipedia.org/wiki/The_Marching_Morons>).
Risk -levels- are a choice and also a trade-off.
The risk is being managed at the VC/investor level, by diversifying investment bets over numerous early ventures.
The death of any one of those isn't a concern for the VC, if the portfolio performance is sufficient. Of course, for the individual venture and employees, that risk is disaggregated.
More rigorous systems practices are seen as an impediment to early growth with any potential problems either something that can be ironed out later, or simply a post-liquidation concern that doesn't factor into the investors' interests at all.
At a mega corporation, you can't take risks which make complete sense for the project you're on independent of uncapped liability for the mothership.
There are misalignments like the one you describe but that's not what I'm talking about.
What’s funny is I seem to have to explain this to senior management anew every 6-7 years. I know they teach it in management school, but it in the real world somehow people fall into the false equivalence when they get promoted. Often they adopt a cartoonish view of things because they can’t get the quantitative signals and everything decision effectively reduces to what I ironically term as anecdotal analytics.
I have this amusing heuristic for risk acceptance which I often use to help people approach decisions: you should kick the decision up to someone with higher authority if your signing authority is less than: risk coefficient times quantified exposure, less mitigation cost where mitigation is within signing authority AND/OR budgeted and authorized spend. I like to view mitigation and opportunity cost/benefit in a similar way so I have some idea of equivalences when evaluating tradeoffs.
It’s not original with me, I must have lifted it from some decades-past HBR article or 60’s rant on quantitative business management.
I could rant on various aspects of risk management application all day, thank goodness I've managed to quit before I really got started sharing...in my experience it’s been very helpful when applied in real world engineering implementations.
Ha, I used to suggest people consider “total failure of business” in exposure quantification...
Upside was that I could definitely know they weren't from the same batch.
Downside was that I had to buy a Seagate, and I don't have good experiences with Seagate since my only Seagate drive had died an early death at the tender age of 3. Turns out that this was very much a downside since the Seagate drive died at the tender age of 16 months.
Anecdotally, I haven't had issues with Seagate but I'm sure it really boils down to which exact drives you're using and what batch they were in.
Some may be worse than others but diversification is the right answer anyway.
This is why my bicycle drivetrain should be a frankenstein combination of parts from different manufacturers?
/s
If we're ever both at the same conference show me a link to this and I'll buy the first round.
haha, pimp your ride!
I find hard drive warranties to be mostly an illusion. It's better now that full disk encryption is becoming better supported and potentially available on personal devices and not just corporate ones managed by IT professionals. However until recently the number of drives I've had in any personal/home system that I would have returned under warranty instead of securely destroying to prevent the risk of data leakage was zero. The number of phones I have ever traded in is similarly zero. It's horribly wasteful but until there are cast iron guarantees that all the private data we keep on these devices is going to be securely deleted it's the only sane policy IMHO (apart from never using these devices for anything remotely sensitive in the first place but that's all but impossible in modern society).
For even more peace of mind, (and only when you can afford it, obviously) try decoupling your disk purchases a bit from when you're going to need them.
When you see a good price or a sale on a particular disk, grab it add it to your own personal "prebought disk pool". When it's time to either replace a disk or spin up a whole new array, now you have the benefit of diversification across time.
The issue here (as it was several years ago with the re-known Seagate 7200.11 issue [0]) is not about the odds of multiple (hardware) failures together (which may actually be a very rare case), in these case it is essentially a software failure, a counter that crashes the on-disk operating system (if we can call it so) be it an overflow of the counter or hitting a certain value.
The chances of having almost simultaneous failures is near to certainty for drives that are booted the same number of times and have been powered for the same number of hours, if the affected counters are related to these events.
[0] Some reference:
https://msfn.org/board/topic/128807-the-solution-for-seagate...
>Root Cause
This condition was introduced by a firmware issue that sets the drive event log to an invalid location causing the drive to become inaccessible.
The firmware issue is that the end boundary of the event log circular buffer (320) was set incorrectly. During Event Log initialization, the boundary condition that defines the end of the Event Log is off by one. During power up, if the Event Log counter is at entry 320, or a multiple of (320 + x*256), and if a particular data pattern (dependent on the type of tester used during the drive manufacturing test process) had been present in the reserved-area system tracks when the drive's reserved-area file system was created during manufacturing, firmware will increment the Event Log pointer past the end of the event log data structure. This error is detected and results in an "Assert Failure", which causes the drive to hang as a failsafe measure. When the drive enters failsafe further update s to the counter become impossible and the condition will remain through subsequent power cycles. The problem only arises if a power cycle initialization occurs when the Event Log is at 320 or some multiple of 256 thereafter. Once a drive is in this state, there is no path to resolve/recover existing failed drives without Seagate technical intervention. For a drive to be susceptible to this issue, it must have both the firmware that contains the issue and have been tested through the specific manufacturing process.
Remind me of my laptop that I bought with 2 SSDs. Not from HP or Dell, but still, I now wonder if I should replace one of them with a more recent SSD and give the other to my son (he currently has an anemic SSD that's too small to install Genshin Impact on).
Just how much difference is enough to be safe is the price question...
Unsurprisingly, 14 years later I still wouldn't recommend Seagate drives to anyone.
May have to go check my up hours on my drives now, I must have a few nearing that sort of write hours
- Unless you periodically do full-drive reads, you may silently accumulate bad blocks across multiple drives in an array. When you finally detect a failed drive, you discover that other drives have also been failing for months.
- A full RAID rebuild is a high-stress event that tries to read every disk block on every drive, as rapidly as possible.
- And finally, some drive batches are just dodgy, and it may not take much to push them over. And if identically dodgy drives are all exposed to exactly the same thermal stress and the same I/O operations, then I guess they might fail close together?
Honestly, RAID arrays only buy you so much reliability. Hardware RAID controllers are another single point of failure. I once lost two drives and a RAID controller all together during Christmas, which was not a fun time.
I do like the modern idea of S3-like storage, where data is replicated over several independent machines, and the controlling software can recover from losing entire servers (or even data centers). It's not a perfect match for everything, but it works great for lots of things.
I used to help manage a large fleet of database servers. We found that blocks could "rot" on the underlying storage, yet if they were read often enough they would be held in memory for months and never re-read from the underlying drive. Until you rebooted!
$ sudo smartctl -a /dev/sda | grep -e Power_On_Hours -e ^ID
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
9 Power_On_Hours 0x0032 098 098 000 Old_age Always - 9743
Just looking at the raw value, it seems to be 9'743 hours in my case...but always check your backups regularly for data that is dear to you!
Protip of the day: that includes things on someone else's server. I remember when Grooveshark went offline from one day to the next and I lost nearly my whole library because I remembered only some artists and had to go through thousands of songs to find which ones I actually liked from them. My browser's localStorage object had the playlists but I didn't use those much. Or when 000webhost cancelled my account because I was using the 100MB(?) to back up some files that were most important to me, rather than for actual webhosting (in my defense, I was 15 at the time), and so when I returned from a holiday with my parents with an actual crashed hard drive, that turned double sour. Backing up things from what they now call the "cloud" is something I learned early, as I have virtually no code I wrote before that summer, only some of the music, only essays with WordArt if they were printed, etc.
If you use Telegram, Spotify, Netflix, maybe you have videos uploaded to YouTube and not have local copies anymore... my recommendation is to have backups of things that are important to you, onto a medium that you own (it's only a disaster-case copy anyway), and for copyrighted content like Spotify/Netflix it would simply be enough to just have a list of songs/videos. Maybe Netflix doesn't go offline from one day to the next, but your account might be hacked, or that friend you share it with might be hacked, or the right set of hard disks fail at their datacenter, etc. GDPR data exports are your friend, particularly when they're automated and you don't have to bother support. (They might also reveal, as in my case, that Spotify knows how often you shower because I then connect it to a particular waterproof bluetooth speaker at wakeup time. Data exports are also fun to browse!)
sudo pacman -S gsmartcontrol
sudo smartctl -a /dev/nvme0n1p4 | grep -e "Power On"
Im at only 2727 hours"Power On Hours: Contains the number of power-on hours. This may not include time that the controller was powered and in a non-operational power state."
The same drive reports only 329 "controller busy time" minutes.
"9 Power_On_Hours 0x0032 001 001 000 Old_age Always - 74233"
12 years old. More than 8 years of run time. It keeps on purring.
Yes, I have redundant backups. I also have a replacement drive ready. I just want to see how far I can take it.
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
9 Power_On_Hours 0x0032 055 055 000 Old_age Always - 39676
I'm 300 hours from 40K, time to buy new SSD? is this real?!Have you applied the appropriate firmware update?
1. Check your backups
Get-PhysicalDisk | Get-StorageReliabilityCounter | Select-Object PowerOnHours
(shell 1) ~# smartctl -a /dev/sda | grep -e Power_On_Hours -e ^ID
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
9 Power_On_Hours_and_Msec 0x0032 000 000 000 Old_age Always - 933932h+27m+33.940s* 100% of 2015-2017, let's add 2 years here
* Aboutish 50% of days since 2018 to 2020
* On and off again (5%?) since then until now.
So it's about 3 years of full use? I'm eyeballing the use here. So it may be close to the numbers that were given, but I'm not sure. Guess I could check the SMART stats to get a precise number and from there decide what to do about it.
Searching a bit it seems it's a well-known bug in "enterprise SSDs"[0, 1] (which my drive certainly isn't) but there aren't any real details about what causes it, other than "a firmaware bug".
[0] https://www.servethehome.com/hpe-issues-hpd7-fix-for-ssds-th...
[1] https://www.anandtech.com/show/15673/dell-hpe-updates-for-40...
The Cisco report turned up in response to a post I'd made of the HN issue on the Fediverse:
None of which gained traction at the time:
OCP (Open Compute Project) has shown that customer-operators can cooperate on open hardware designs, successfully influencing enterprise hardware supply chains. Commercial DPUs and SmartNICs were preceded by a decade of open hardware and research by the NetFPGA project (https://netfpga.org). Why not DiskFPGA?
2017 OpenSSD overview, based on Xilinx: https://github.com/Cosmos-OpenSSD/Cosmos-plus-OpenSSD/blob/m...
2022 status, http://www.openssd-project.org/
> OpenSSD platforms are still being actively used in many academic institutions. As of June 2022, we have renewed the homepage hoping that this site will be a forum to share various simulators, tools, traces, etc. not only for the conventional SSDs but also for the upcoming storage devices such as KVSSD, ZNS SSD, and Computational Storage (CSX). This site is being maintained by Systems Software and Architecture Lab. at Seoul National University as a part of the SW STAR Lab. project.
Edit: this obviously ignores any troubles one would have sourcing the ICs (such as possible NDAs)
[1] https://www.mouser.ca/datasheet/2/671/micron_technology_mict...
So using that flash, you could build something that is recognizably an SSD. But it would be almost entirely useless: too expensive and too small and slow for production use, and too far removed from the current state of the art to serve as a research platform for the most important challenges the SSD industry has been dealing with for the past several generations (error correction strategies for TLC/QLC, and SLC caching).
- I just picked one at random. I'm sure the bleeding edge is harder to get and datasheets are harder to get, but I wasn't trying to find the newest or best.
- the specific subthread here is about the diy-ishness of ssds vs. spinning rust, where the difficulties are of a fundamentally different kind. I feel like it goes without saying that a home built ssd is not going to perform to the level of mass production devices, the question was just can you.
Also, with these multiple level flash technologies (quad level is current tech, triple is still used in some SSD/NVMe) the read, write, and ECC algorithms are non-trivial to the point where last I checked even mainline Linux's raw flash driver support won't do anything beyond single level cell flashes (and very few new embedded designs are choosing raw parallel NAND flash, instead opting for things like eMMC or UFS which have built-in controllers to handle this).
The comparison here is that no matter how much you hunt on digikey you won't find a disk platter or drive head or any of the other precision machined parts that go into a hard drive (never mind putting them together and keeping dust out etc).
But they all rely on leaked manufacturer firmware production tools, not open source firmware.
Luckily I had Time Machine backups of my iOS backups and I managed to avoid losing too much data.
As a sidenote it seems like Apple has pretty much neglected their offline backup and syncing workflow to drive more people to just pay for iCloud storage. Half the time my iPhone takes hours just to get detected by the mac when plugged in.
I don't have issues with my computer (PC or Mac) detecting my iPhone. Generally need to make sure iPhone is unlocked after plugging it in. What is tough is the large size of my iPhone (X gb) and how small my Mac's HD is (2X gb).
You actually bring up another issue. There is no obvious way to backup iPhone locally to an external hard drive. So either pay the mac SSD storage tax or the icloud tax.
Again, depends on usecase but then it becomes integral to your existing workflow instead of an addendum that you end up forgetting to do
The whole purpose is to make the failing of one be effectively extremely noisy and irritating
It's like what I do with raid. I have a script that will shut the machine down on drive failure and then will use dialog(1) to say something like "hey bozo replace the fucking drive first" when you boot it up and then it will shutdown again and be unusable.
Make the complaining show stopping, loud, rude, and disruptive. Because if the next one fails you're screwed
Of course your SSD might have some other firmware bug that would eat your data, all you can do is search for the model number and see if the manufacturer has issued any notices/firmware updates.
That’s just your presumptive opinion, right?
Edit: sorry, probably put that offensively. mikiem said about the HN drives: “These were made by SanDisk (SanDisk Optimus Lightning II) and the number of hours is between 39,984 and 40,032...” - https://news.ycombinator.com/item?id=32031428 Without knowing parts of a codebase are shared between SanDisk devices, it is hard to say that enterprise SAS devices have absolutely no code shared with consumer devices. So just the commenter’s opinion unless the commenter has knowledge of writing Sandisk firmware. “HPE and Dell both used the same upstream supplier (believed to be SanDisk) for SSD controllers” https://www.anandtech.com/show/15673/dell-hpe-updates-for-40...
Even if the code containing this bug was shared between consumer and enterprise drives, it's not reasonable to assume that it would take SanDisk multiple years to check whether their consumer drives are also affected. The lack of a follow-up report from SanDisk is good evidence that their other products are not affected.
However I dislike a black and white “This will not” absolute fact statement: even if based on reasonable assumptions which is what appears to be the case versus detailed knowledge.
Most laptops don’t run their SSDs 24/7, and unless a manufacturer’s error affects a lot of consumers, we often don’t find out the cause of consumer equipment errors in my experience.
If the OP has a laptop older than 2020, with an SSD with a crucial chipset (especially if SATA), and they leave it on most of the time, then maybe check SMART hours.
That drive launched in 2011 though so there probably aren't that many still in active use which still haven't reached ~7 months of uptime.
That problem became known a decade ago, so it's somewhat surprising to see such a similar bug now.
This new one is worse because the drive cannot be used after reaching the magic number of hours. In the Crucial M4 case the firmware could be updated even after the bug struck.
Had to rebuild from HDD backups, down for a week. I still have nightmares.
https://en.wikipedia.org/wiki/OpenBIOS
>OpenBIOS is a project aiming to provide free and open source implementations of Open Firmware. It is also the name of such an implementation.
>Most of the implementations provided by OpenBIOS rely on additional lower-level firmware for hardware initialization, such as coreboot or Das U-Boot.
https://en.wikipedia.org/wiki/Open_Firmware
>Open Firmware is a standard defining the interfaces of a computer firmware system, formerly endorsed by the Institute of Electrical and Electronics Engineers (IEEE). It originated at Sun Microsystems, where it was known as OpenBoot, and has been used by vendors including Sun, Apple, IBM and ARM. Open Firmware allows the system to load platform-independent drivers directly from a PCI device, improving compatibility.
Of course, I am also entertaining the possibility that no one thought they would be in use for this long, which would certainly be evidence of planned obsolescence.
No sane SSD manufacturer would do such thing on purpose. You do it and you loose business, that's it.
The simplest explanation is that somebody made an honest engineering mistake.
If we do the opposite, as you say, and assume everything is an honest mistake, that puts pressure on the single customer to prove that the organization with a huge marketing budget is doing something wrong. In this situation, the worst thing that happens is we all get taken advantage of.
Our collective distrust is the only power we have against massive marketing/PR budgets. It doesn't have to be angry, or sour, or cranky, we just collectively need to not take their word until we have a reason to do so.
'This insulation prematurely disintegrates under normal use causing the wires it is designed to protect and insulate, to short causing many problems.'
http://www.mercedesdefects.com/2008/04/wire-harness-defect.h...
Other than "intentionally" (which we cannot know and makes no difference to whether you lose your data or not) that is literally what these SSDs are doing, and no SSD brand has been destroyed over it.
Also, enterprise drives get firmware updates (regardless of spinning or not), and this firmware is automatically applied via RAID controller, so it could be remedied easily before it got this big if it's an actual error.
Unless you burn through your SSDs, you're very unlikely to hit this event.
When these servers' continue to be used and disks all start to fail at the same time, this will obviously stink.
The bathtub curve is not like this. You can feel that.
Now more than ever, five year old server hardware isn't that far behind the curve unless you're on the bleeding edge. I've been looking for bottom of the barrel hosting lately, and there's lots of dedicated servers available with 10+ year old cpus, and probably most of the rest of the machine is a similar age.
>>> 2**57/10**9/3600
40031.996687737745Though that might cause a divide by zero?
What could cause unexpected behavior at 57 bits?
Perhaps storing fractions of an hour, like incrementing it every 1/16th of an hour and calculating a relative rate of change, causing a divide by zero?
Engineer A: Gee, I need to store a few flags with each block, but there's nowhere to put them. Ah! We're storing timestamps as 64-bit microseconds. I can borrow a few of those bits and there'll still be enough to go for thousands of years without overflowing.
Engineer B: Gee, our SSDs are getting so fast, soon we'll be able to hit 1M writes/sec. But we're storing timestamps as microseconds. How can we generate unique timestamps for each write? Ah! I'll switch to nanoseconds. It's a good thing we have plenty of space in this 64-bit int.
BOOM!
52 is notable as 2^4 + 2^3 = 24, 24 + 24 = 48, 48 + 2^2 = 52. But 57?
"The fault fixed by the Dell EMC firmware concerns an Assert function which had a bad check to validate the value of a circular buffer’s index value. Instead of checking the maximum value as N, it checked for N-1. The fix corrects the assert check to use the maximum value as N."
https://www.anandtech.com/show/15673/dell-hpe-updates-for-40...
Why the MAX value would be in an circular buffer, or what was being stored in N-1? No idea.
The infamous Seagate firmware bug was due to the same thing.
I presume you find a lot of circular buffers in SSD firmware, for wear-leveling reasons. Samsung's NILFS and NILFS2 are structured as circular buffer append-only logs, at least partly to avoid trusting the firmware wear-leveling.
Any time I see these sorts of issues (odometers that kill themselves, for example) I think of smaller units at higher bit depths. That's not the only way to get to this kind of concern, but it's a way that pretty-darn-competent engineers can leave ticking time bombs due to estimation failures.
Always check types for overflow and/or precision loss. Always.
https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na...
https://support.hpe.com/hpesc/public/docDisplay?docLocale=en...
> Are any consumer drives affected by the same issue?
As far as I know it doesn't affect consumer drives, but I wouldn't be surprised if some have the same defective firmware.
edit: apparently sandisk: http://forum.hddguru.com/viewtopic.php?f=3&t=39964 - also a clue about the magic 40k hour significance: "the SSD alters its performance in some way as it approaches end of life. This appears to shine some light on the reason for a trigger at 40K Power On Hours."
https://www.anandtech.com/show/15673/dell-hpe-updates-for-40...
I have a Samsung EVO and OCZ SSD. Would these be affected too? Perhaps some shared component?
Cisco has written "Industry-wide" here which is confusing
While the concept of Chia was interesting at the time and also reminded me of the "smart fridges of silicon valley (the show)", filling up gigabytes with trash data to prove a technical point made me lose interest. Just glad I didn't invest more.
Considering this is a firmware "bug" that bricks the drive, due to a misplaced index, not a physical wear issue, it appears to be a 4.5-year planned obsolescence feature.
"...Once a [SSD] drive has surpassed the 43,800 hour mark (5 years), it may no longer be classed as in "perfect condition" "
And that SSD generally has 5 year life expectancy.
So with this bug, should we simply think of it just to become 40,000 hour hard life-time limit? Well, it's 10% less than by design.
I'm just not sure how realistic will it be to obtain SSD firmware updates given that it's "an industry wide firmware index bug".
How could I even know if a particular SSD has an affected firmware?
Especially not electronics. And certainly not more advanced semiconductors.
The rule-of-thumb is the lifespan of planar CMOS processes in years is proportional to the "node size" in nanometers. So we are right now at 2-10 years. Some specific ICs or specific applications or specific designs can be bigger or smaller than this but it's the average.
If you've ever worked on "antique" electronics, this is no surprise.
I didn't have `smartctl` installed on my Mac, so here's how to install it via `brew`:
brew install smartmontools
My SSD is `disk0` (check `Disk Utility.app`).Running `smartctl` for my `disk0`:
sudo smartctl -a /dev/disk0 | grep -e "Power On Hours"
Power On Hours: 6,334
So I have only ~15% of those 40k hours used (6,334/40,000).In that case, I would expect them to be rounding down, not up, and to at least mention the exact number somewhere.
> but re HN they’re counting clock time not power on time. Could have easily been turned off total 10 days for maintenance etc.
I seriously doubt HN has been down for 10 days over the last five years. As discussed in the other thread I linked, this has been one of the longest HN outages.
My SSD's were scragged. my HDD's were just fine. I guess it's time to figure out how to get a realistic write-through cache setup going, because from now on, if it ain't on spinnin' rust it ain't hard enough yet.
CPU/GPU and RAM lived, but the corruption of the drives (even if the data was largely recoverable) and rendering of them as inviable to further writes really took me by surprise.
That combined with the way the HDD just did not care one lick despite apparently, just illustrated for me a difference in tolerance to operating conditions that I'd not had the chance to witness first hand yet.
Just figured I'd share while we were talking about SSD weirdness and firmware nonsense.
Edit: If it's a problem in Cisco's upstream vendor then it could affect others, but probably still just enterprise stuff.
> HPE is one of the SSD OEMs affected by it: ...