CPU reliability – Linus Torvalds (2007)
yarchive.net
yarchive.net
The only times I've even heard about failing CPUs has been if they've been overclocked or insufficiently cooled(add in overvolting, and you get both :)) or physical damage during mounting/unmounting or otherwise handling hardware. And even then the failure has usually been elsewhere than the CPU itself.
Of course I am not saying it'd be unheard of, but for me frankly, right now it is.
That sounds like a huge cost to bear. Looking at e.g. a Haswell die photo [1], there are just four physical cores present for a four-core part. With that die area per core, you would take a ~15-20% area hit (that translates to 15-20% cost) just to have a spare core in case one failed some years later.
I have heard of manufacturers selling otherwise "defective" parts where a core or cache slice has a defect by relabeling as a part with fewer cores. But that's a manufacture-time decision, not a dynamic reconfiguration in the field.
[1] http://cdn2.wccftech.com/wp-content/uploads/2013/05/Intel-Ha...
It seems like the GP might be suggesting that a quad-core CPU will swap in another core when one dies. That doesn't happen. But the binning process allows them to still sell slightly defective silicon with disabled parts (cores, cache), which saves money.
On a related note, a lot of GPUs actually do have a few dozen execution units that are disabled by default and can be swapped in after stress testing at the factory. I believe some can even do that in the wild, but I could be wrong.
https://en.wikipedia.org/wiki/PlayStation_3#Technical_specif...
The cell processor had a main PowerPC core and eight floating-point SIMD co-processors called synergistic processing elements (SPE). For the PS3, one of the eight SPEs was disabled and another was reserved for the operating system, leaving the other six for developers.
Failures were more prominent in memory...but they did happen. We also sent equipment through environmental testing that would force failures. I don't recall of hearing of any CPU failures. Although most of our equipment was DSP & FPGA based but there were some tiny 'lil CPU's there.
Does this mean it didn't boot into Linux?
It would actually run for a day or two, serving up a modest PHP load, before crapping out.
It's entirely possible that they overheated, but if they did it was due to poor cooling; we did not overvolt or overclock these machines.
In summary, some data centers are run hotter than recommended, which leads to a lot of mostly ignored domain resolution errors, which leads to a security risk.
Or how you notice how much more frequent your model of car is on the roads?
Another possible factor is that I find I pay more attention to things I've recently learned about more closely when I come across them, but ignore them when I either know them very well or don't know them at all. I don't know if this is true for anyone else.
The first would include the Kardasians for example. But just because they get mentioned a lot, you don't feel anything strange about hearing about them a lot.
Whereas the very essense of the phenomenon is that encoutering something multiple times seems strange to you.
Several. But I'm lucky in that I worked for NetApp for 5 years which have several million NetApp filers in the field that were all calling home when they had issues, and Google which has a very large number of CPUs all around the planet doing their bidding. With visibility into a population like that you see faults that are once in a billion happen about once a month :-).
Two general kinds of failures though, the more common one is a system machine check (the internal logic detected the fault condition and put the CPU into the machine check state) which happens when 3 or more bits go sideways in the various RAM inside the mesh of execution units. Nominally ECC protected it can detect but not correct multi-bit errors. Power it off, power it on, and restart from a known condition and it's good as new.
The more rare occurrence is that something in the CPU fails which results in the CPU not coming out of RESET even, or immediately going into a machine check state. When you find those Intel often wants you to send it back to them so they can do failure analysis on it. The most common root cause analysis for those is some moderate electrostatic damage which took a while to finally finish the process of failing.
Some of the more interesting papers at the ISSCC are sometimes on lifetime expectancy of small geometry transistors. They are a lot more susceptible to damage and disruption due to cosmic rays and other environmental agents.
Microsoft released research that shows enthusiast overclocking has a clear increase in hardware faults.
[1] http://www.extremetech.com/gaming/131739-microsoft-analyzes-... [2] http://research.microsoft.com/pubs/144888/eurosys84-nighting...
[14865975.000023] Machine check events logged
MCE 0
CPU 0 BANK 2
ADDR 1438280
TIME 1384859595 Wed Nov 20 00:13:15 2013
STATUS d40040000000011a MCGSTATUS 0
MCGCAP 104 APICID 0 SOCKETID 0
CPUID Vendor AMD Family 6 Model 8
This will, of course, be getting replaced shortly, preferably before it does any real damage. Given it's a 2002-era chip, 11 years of service isn't exactly terrible.
I've seen quite a few UltraSPARC chips (especially IIIs) go over the years at work, and often had a shit of a time trying to get Sun to accept them as faulty and replace them.
Those were correctable errors, prime95 or memtest did not detect anything.
And the failed part is either needed only on bootup or when the current gets going it doesn't stop until it's powered off again
Which wasn't a real problem for 'some' day-to-day use. This was in the mid to later 00's, so https wasn't quiet everywhere yet.
The problem manifested slowly. When ever I'd connect to HTTPS, my browser would crash. One sound card would phone home for an update, and my computer would crash. Randomly certain games would crash when ever anti-cheat software attempted to run.
It was just odd, and took a few days of hunting to find out what was actually going wrong.
In the real world this doesn't happen very often and there are techniques to mitigate it when it does (usually at a performance or latency cost) - core CPUs are probably safe, they're all one clock but display controllers, networking, anything that touches the real world has to synchronize with it.
For example I was involved with designing a PC graphics chip in the mid '90s - we did the calculations around metastability (we had 3 clock domains and 2 crossings), we calculated that our chip would suffer from metastability (might be as simple as a burble on one frame of a screen, or a complete breakdown) about once every 70 years - we decided we could live with that as they were running on Win95 systems - no one would ever notice
Everyone who designs real world systems should be doing that math - more than one clock domain is a no no in life support rated systems - your pacemaker for example
You can design to be metastablity tolerant - use high-gain, high clk->Q flops as synchronizers, uses multiple synchronizers in a row (trading latency for reliability), you can do things to reduce frequencies (run multiple synchronizers in parallel, synchronize edges rather than absolute values etc), but in the end if you're synchronizing an asynchronous event you can't engineer metastability out of your design - you just have to make it "good enough" for some value of good enough that will keep marketing and legal happy.
It's our dirty little secret (by 'our' I mean the whole industry)
As a consumer it's hard enough to keep up with what's reliable in hard drives. Keeping the manufacturers honest with good stats for the most common parts would be great.
https://static.googleusercontent.com/external_content/untrus...
The backblaze guys also have a big data set:
http://blog.backblaze.com/2013/11/12/how-long-do-disk-drives...
Table 3 suggests that there are data sets that include all components (CPU, memory, power supplies, etc.).
Any large company that makes things employs a bunch of reliability engineers, who are usually EE's or ME's who make Weibull plots and bathtub curves all day (to set the warranty duration, mostly). These guys have all the data you could ever want on this topic, but they're not sharing. Especially at Intel.
I know Intel has been working for some time on the idea of high temperature data centres, this will impact the MTBF of all components but you can always calculate the cost of the losses vs the cost of the cooling: http://www.datacenterdynamics.com/focus/archive/2012/08/inte...
Found it here, which also goes into some testing the Guild Wars guys did on their population of gamer PCs: http://www.codeofhonor.com/blog/whose-bug-is-this-anyway (scroll down to "Your computer is broken", around 1% of the systems they tested failed a CPU-to-RAM consistency stress test)
Both of them indicate intermittently defective components in running systems are way more common than anybody assumes.
But when they say "Memory Error", even though it's something detected/corrected by ECC I'm not sure we can say 'the memory is defective'
It may be a combination of the conditions of power/load/data/time since last refresh and variance between modules.
Since Google appears not to show all the data they have, we probably are not going to get that from them though :/
One needs redundant hardware to provide certain guarantees about the service being up. This means load balancers, multiple CPUs running the same code in parallel and comparing results, running on separate power buses, different data centers, different parts of the world.
> different parts of the world.
Still takes just one asteroid.After a certain number of 9s you just have to smile, nod, and truncate the number.
https://aws.amazon.com/s3/faqs/#How_durable_is_Amazon_S3
I suppose that if they ever lose an object then they can say "well we warned you that you might lose an object every hundred million years."
On the other hand, it's unlikely that many people would be around to ask for a reclamation. It's more likely that Amazon would go out of business before anything like this happened, anyway.
At one of my last jobs someone put the following ticket in the bug tracker "Following the end of the world on 12/21/2012 the system Blah will stop working"
After the date it was closed with a "As the end of the world didn't happen this ticket is no longer needed"
"Cycles, Cells and Platters: An Empirical Analysis of Hardware Failures on a Million Consumer PCs"
http://research.microsoft.com/apps/pubs/default.aspx?id=1448...
Also, I'd imagine that space craft components are of an entirely different category of components that the off the shelf computing variety.
It's probably not fair to compare the MTBF of specialised hardware to the $35 CPU I bought at the retailer down the street either, the RAD750 processors in Curiosity cost almost a quarter of a million dollars each.
http://en.wikipedia.org/wiki/Comparison_of_embedded_computer...
http://en.wikipedia.org/wiki/Curiosity_rover#Specifications
http://en.wikipedia.org/wiki/Radiation_hardening#Radiation-h...
Though that said, Voyager is still happy running on it's 8064 words of 16 bit RAM, which is something.
"The frequency of the heartbeat, roughly 30 times per minute, caused concern [176] that the CCS would be worn out processing it. Mission Operations estimated that the CCS would have to be active 3% to 4% of the time, whereas the Viking Orbiter computer had trouble if it was more than 0.2% active15. As it turns out, this worry was unwarranted."
They are using DMA a lot; instruments write to memory, occasionally the CPU is turned on and picks up the new values. Also they had to manage with the fact that memory is degrading, so the system needs to adapt to working with less memory. The bus is 16 bits wide, but actually they are processing 4 bits at a time, so addition takes 4 cycles. CPU registers are stored in RAM, so probably they can reassign them if a memory cell fails.
Parts of the system were reused from the Viking mission. Also they where reprogramming the system in flight during the eighties ! That's the reason why they could start the the mission, even without having the full software on board, the mission was extended thanks to reprogramming. Just for the Jupiter visit they had 18 software updates, think about that next time that a software update breaks something on your system.
Also its all a distributed system with several CPU's, and some elements of redundancy, awesome tech. I guess one day alien hackers will have fun with reverse engineering this system.
Now this one has to work 24/7 in a hostile environment; has to be hidden; has to deal with enormous quantities of data and it costs a lot to replace/repair so it must be very reliable.
What is driving technological progress? Instead of a space program, we now have political control of the Internet as driving forces. I guess that's what they mean when they say that civilization is turning inwards ;-)
Yes, in many areas the NSA and Google are pushing the envelope; long term data storage; map reduce of large data sets; AI, you name it, they have it.
Weren't they using hardware in submarines anyway?
The equipment doesn't have to actually duplicate L2 frames. It just uses standard fiber repeaters (already a common component in undersea cables) to get it back to a more friendly environment where they can actually decode and process it.
Apart from electron migration issues and failures by excess (voltage/temperature), they're pretty long lasting
Much easier to have a failure because of something else: capacitors failing, oxidation or mechanical failure (for example, because of thermal expansion/contraction)
I've seen people complaining about a dead CPU but I can't find it right now
I design integrated circuits and one of the constraints in selecting the width of wires is to make sure that the maximum current density is below the electromigration threshold.
Overall though, given with how many computers I've worked with, CPU failures still seem rarer than Memory, Disk, Mobo, or Graphics failures. Of course it ends up being the CPU in _my_ computer that fails -.-
How long did it work for before it died?
There is still some variance in silicon, so yours may have had a defect that manifested itself after some time, I'm not sure they evaluate returned defective chips to see what happened (and if this is public info)
Also, the packaging is extremely complex and prone to the same kind of defects as other PCBs in the system.
Disk failures are very common, followed by much rarer RAM chips and motherboards failures.
I suspect server chips are rated for 10-15 years average lifespan
Does that mean a CPU/RAM/GPU will not perform as well as when it's brand new ?
RAM yes, PROMs yes, CMOS batteries yes, PSUs yes, drives yes.
They're probably the most reliable bit of a computer.