An Empirical Analysis of Hardware Failures on a Million Consumer PCs
research.microsoft.com
research.microsoft.com
The only thing I would have wanted to see but didn't in this analysis is how failure rates vary for different types of disk subsystem -- specifically, traditional hard drives versus the newer solid-state devices. I suspect, but don't know for sure, that the latter have much, much lower real-world failure rates in the first 30 days of total accumulated CPU time (TACT).
The authors openly suggest that the sharp difference in failure rates between desktop and laptop machines may be due in part to their disk subsystems: "Laptops are between 25% and 60% less likely than desktop machines to crash from a hardware fault over the first 30 days of observed TACT. We hypothesize that the durability features built into laptops (such as motion-robust hard drives) make these machines more robust to failures in general." Alas, the authors don't delve any further into it.
I'd like to see hard data comparing the real-world failure rates of both desktops and laptops using traditional versus solid-state disk subsystems.
My take: so far, no one has sufficient consistent-across-the-board data at the moment to reach a conclusion about the matter, but the anecdotal evidence presented in that article suggests that Intel SSDs probably have lower failure rates than most traditional and solid-state alternatives. I will keep that in mind next time I buy an SSD.
There are also numbers for hard drives: http://forums.anandtech.com/showthread.php?t=2147063
Basically, Intel SSDs from a year ago are more reliable than all hard drives. And SSDs in general are more reliable than any 2TB hard drive.
The data isn't ideal, but it's better than anecdotes. Return rates should correlate pretty well with failure rates. If anything, return rates should favor hard drives, since people are less likely to return a faulty cheap hard drive than a faulty expensive SSD.
In short, the available return-rate data is too noisy and inconsistent to be a good proxy for failure rates.
[1] http://www.tomshardware.com/reviews/ssd-reliability-failure-...
Lack of moving parts is great, but flash allows a finite number of write cycles.
I think Intel released some information on the reliability of their SSDs a few years ago but that was likely because they knew they were doing best and their enterprise customers are very interested in that for their data-centre rollouts.
The very limited information I've seen suggests to me that a few years ago SSDs had a much higher failure rate than HDDs (double in the first 6 months) but that has been falling very quickly with each new generation of SSDs (and as the profit margin grows and manufacturers have to work on reputation to justify the markup).
Some people over on xtremesystems have done Endurance testing, and the 64 GB m4 took over 700TB of writing to for failure to occur, and 172 TB to reduce the MWI to 0.
In a little over a year I have only written 4.1TB to my SSD in my desktop. Write cycles are very unlikely to run out for me before I replace the drive.
http://www.xtremesystems.org/forums/showthread.php?271063-SS...
It's talking specifically about disk failures though, not comparing whole systems
I understand the reasons (not alienating your hardware vendors), but will there ever be a research group who will disclose vendor names? Heck, I would pay for this information.
* a machine that crashed once is 100 times more likely to crash again; the more it crashes, the more it's prone to fail again.
* overclocking significantly reduces reliability. One CPU vendor (AMD or Intel, but unspecified) is much worse in this regard, too.
* conversely, underclocking improves reliability.
* branded computers are more reliable than beige boxes.
* laptops are more reliable than desktops.
This is meaningless.
The real question is how those compare to beige box that uses a decent parts. But Microsoft definitely has interest in helping manufacturers of brand computers because piracy is more prevalent in beige boxes.
Disclaimer: I am a researcher in a corporate lab.
Even doctors show this kind of biases when advising people on choosing treatments(which is a much bigger moral issue).
1. Their data collection methods are biased against white boxes. Given the large sample size and the method of retrieving samples - automatically generated crash reports from users - I find this unlikely. They cover this point in section 3.2.1.
2. Their statistical analysis is flawed. I see no issues with it, nor did the reviewers. (Otherwise it wouldn't have been accepted.)
3. They lied. I am most skeptical on this one.
It's disingenuous to gesture at researchers and allege bias based on their employer without actually saying how they are biased. Doing so is not valid skepticism, but prejudice.
Do you think they haven't known that ?
This kind of remarks of the incompleteness of research done is common in many research papers. And they definitely contribute to readers not getting the wrong impression.
The authors stated a conclusion, but did not speculate on the cause behind the conclusion. I see no bias in them not calling attention to the fact that they have not studied the cause - that is self-evident.
It's also possible that it's getting more difficult to accurately spec systems, to enforce vendor quality (Dell gets a bad batch of drives, they can 1) detect it and 2) tell the vendor to stuff it, Ahmed's Boxez'R'Us may not have that leverage or depth of experience), and to do burn-in testing of their own systems.
That said, I've had good and bad experiences with big-name and white box vendors alike.
For example, some have speculated that "Gamer RAM" with mean looking heatsinks is actually poorer quality stuff that requires additional cooling to work correctly.
I assume more "beige boxes" are either bought together from cheap components and been assembled by the users themselves or are generally cheap noname buys assembled by who-knows in the shop or maybe they have been modified and/or over-clocked by enthusiasts thus making it more likely to fail - whereas the typical users who buy "brand name" PCs or laptops are not going to mess with them and they can rely on at least SOME standardized quality assurance and control.
A large mfg, I would imagine, would test a configuration repeatedly before making it available, and then, once approved, the individual systems would go through burn-in, probably with more rigor than beige boxes. So even beige boxes with pricier (but unproven configurations) might suffer from grater failure rates. In addition, large mfgs might be able to demand better "lots" from their parts mfgs/oems.
Just a thought.
Essentially you the consumer are doing the burn-in. Its cheaper for Dell to replace failed machines. The cost to burn-in (and the time!) is large.
Everything we've learned from experience, surveys, and PC World magazines has showed the opposite. Heat kills hardware and laptops have their hardware packed together so closely that it generates lots of heat. Back then I remember reading something like 1 in 4 laptops fail in the first 3 years. Which was very believable, at the time I was in collage for game design & development. All 80 guys in our class had laptops from HP (with get this... Pentium 4s in them). Those laptops had a LOT of problems. They were basically portable heaters.
So I guess laptops now have either much better cooling, much cooler CPUs or a combination. OR PCs are just terribly cooled.
"we only count failures within the first 30 days of TACT (total accumulated computing time), for machines with at least 30 days of TACT."
We hypothesize that the durability features built into laptops (such as motion-robust hard drives) make these machines more robust to failures in general. Perhaps the physical environment of a home or of- fice is not much gentler to a computer as the difference in engineering warrants.
If a machine crashes so severely that a crash report is not generated, than those reports will not be present in our data. Therefore, our analysis can be considered conservative...
"Conservative" here means "it underestimates the crash rate by some unknowable amount."
If you drop your laptop down a flight of stairs, it will likely develop some hardware problems. But for one thing, not all hardware failures will cause software crashes and crash reports. For another thing, you're likely to just replace the drop-kicked laptop, and the hardware failures will never appear in logs like these.
One question I still have is whether the switching of CPU frequencies has any effect, or if it is only the average speed that correlates to the reliability. Anecdotal evidence suggests that this is the case, but it could be an area for further research.
- Power supply - Hard Drive
Ranking near these are sleeve bearing fans with ball bearing fans a close second--for CPU and case cooling.
In rough order, I would say that the following is my estimation of other common component failure sources:
- Removable Drives (floppy, optical, etc.) - Video Card (if separate) - Motherboard - RAM - CPU
These are for our corporate PCs which have been Compaq, Dell, IBM, HP, Lenovo and a few other brands.
For our brand name and whitebox server hardware, it's pretty much the same... if something is going to fail, it's going to be a power supply or a hard drive. In fact, I don't ever remember a single server motherboard, RAID controller, RAM stick, CPU or other component ever going bad in a server.
I wonder why they would leave out statistics relating to power supplies when they are, in my experience, the component with the greatest failure rate.
Most important thing in a PC I have found for reliability above everything else is a good PSU, realy does make a difference on the hardware side as you give your kit cleaner power. Add UPS/surge protector and you can double the lifetime of kit. Least from experience I've had it has been noticable.
Here's another copy of the paper:
http://eurosys2011.cs.uni-salzburg.at/pdf/eurosys2011-nighti...
"The table shows that CPUs from Vendor A are nearly 20x as likely to crash a machine during the 8 month observation period when they are overclocked, and CPUs from Vendor B are over 4x as likely"
Obviously it's 5 times difference in probability to have unstable system if overclocked between Intel and AMD but they don't say which one is better. Anybody knows?
Right now, Intel's fastest desktop chip is an i7 990X, which is $1,029 on Newegg. AMD's is an FX-8150, at $199.
Intel prices pretty fairly against AMD on the price/performance curve where AMD has a competitor, eg, the Core i5 3550 at $209 generally outperforms AMD's fastest chip. Pricing then soars off into the sky.
Which is to say, if AMD bumps the speed of their CPUs, they'll release a faster product, compete better against Intel (until Intel reacts), and make more money. If Intel bumps its speeds, they're competing against nobody but themselves, so they usually don't bother.
Therefore, you usually see Intel quite conservatively binning their chips, and they have a lot of headroom. It's not unusual to have an AMD chip that can't go 200mhz faster on air cooling, and to have an Intel chip that can go well in excess of 1ghz faster. So all else equal, bumping Intel chips is less likely to be an issue.
Now, there are some other factors at play here. Firstly, hardcore overclockers pretty much only buy Intel chips. Also, people that are serious often turn up the speed until right before the moment at which the chip starts getting SuperPI errors (i.e. errors at nearly maximum load).
But my totally uninformed gut feeling is that the majority of overclocks aren't done that seriously (if they are, it could actually reverse this analysis).
If we assume that, turning the knob up on an Intel chip without sophistication is much less likely to end badly. The crashes in their paper are not very frequent on average (months between), which isn't necessarily bad enough to revert to old CPU speeds even if the user knows that's what is going on.
> However, in this paper, we do not show a breakdown of drives per manufacturer, model, or vintage due to the proprietary nature of these data.
The much larger weight and volume of desktops would seem to make them easier to cool.
Least I have found desktop's with a good quality PSU and UPS noticable more reliable than those without.
I've heard this theory before but never seen it substantiated. Do you have a link to any research or articles that explains how this works?
But, that doesn't actually answer your question, which I think is "Barely." I feel silly preparing PDFs for publication when I know that most people will read it on their computer, not print it out. Many conferences no longer even have an actual, physical copy of the proceedings, instead just giving out USB sticks with all of the PDFs. (Which is what we want anyway.)
I think it would be fantastic if there was a standard HTML5 template that researchers could use to publish their papers. There are Latex-to-HTML compilers, but I've never been impressed with the results. I think people outside of academia would be more likely to read our papers if they were in HTML rather and PDF.
One reason for PDFs is as you mentioned, latex to HTML results are typically poor. Diagrams are another difficulty that don't have an easy HTML solution. Other reasons I prefer PDFs are: though I never print, I often save papers to disk since I don't always find the paper when I go searching the second time (especially if it is months or even years after), there is a real benefit to being able to read a paper offline - I don't always have a connection to the net when I want to read and lastly, if you have an ereader such as the kindle, pdfs render well on them.
Buying from a lower bin, you're getting a crappier processor, which might give you less latitude to save heat. This would probably be worth measuring and writing an article about. Also, I tend to buy lower clocked processors as it is.