All the other bits in a computer I've seen fail.
All the other bits in a computer I've seen fail.
The component category I've never seen fail is CPU.
Overall I have made very good experiences with long hardware warranties. Soldered RAM on my ThinkPad died once and I got an entire new mainboard after almost 3 years (it died 2 months before warranty expired.) New one still works today but the hardware is just too old.
I’d say on average DC PSUs are vastly overspeced for most machines. Most machines we work with are not sitting at peak load 24x7, and might be drawing half or so of their max rated power draw. Plus every server is redundant - so the modules themselves are operating at around 20-40% of full rated capacity most of their life. The only time they ever spin up past 50% is during power maintenance or a power failure of some sort that brings one of the redundant power sides down.
They do probably operate in slightly hotter conditions on average, but in much more stabilized temperature environments. Plus of course 24x7x365. But that goes for all components.
I’d say other than fans, failure rates in the PC world vs datacenter world are relatively comparable it terms of what fails most. Dust and temp swings are the large variables the average PC needs to contend with vs the datacenter.
Storage of course is the one that stands out the most - everything else is kind of a rounding error. Here though, datacenter work really does beat on the hardware more than most consumer gear so that’s expected. For every 100 disk swaps we do, I figure we do about 2 or 3 power supply modules. And maybe 1 fan swap.
The CPU can fail if the heatsink falls off! That is something that happens (I've had one come off in my hand that was being held on by nothing but hopes and dreams, fortunately not while the CPU was powered on). More likely in a desktop tower form factor.
Somewhat recently, Firefox implemented memory testing in their crash reporter and found 10% of reports were due to bit-flips[0].
If it weren't for reading that HN thread, I would never have being prompted to check and find the faulty memory in my laptop. It would have gone undetected and continued to corrupt whatever was stored in the affected cells. I did run memtest86 not long after I bought the laptop, so sometime between running that first memtest86 and the 12 months after, it developed the fault.
Non-ECC memory sometimes fails and there's a good chance you'll never know.
Memory seems to fail pretty commonly overall, less common than harddisks but more common than motherboards.
The issue is that I'm used to using ECC ram, which fails loud when it's actually bad.. consumer memory won't tell you unless you can't boot.. and everyone disables the startup memory testing too.. (in fact, I think it's disabled by default for the last 10 years because people want to boot quickly).
I used to game on it and I think the design just couldn't handle the cpu and dgpu being active for long periods, it would get extremely hot in the area close to where the memory was.
Granted, I've been building my own computer for 30 years and I've only seen it happen ONCE.
But I've seen plenty of other failures:
- Two hard drives -- One got dropped on the floor, so it was no surprise, the other was showing degraded performance and SMART showed some scary numbers and I was able to replace it before I lost data.
- One GPU, replaced under warranty.
- One AIO water cooler, replaced under warranty. Somehow, the water vanished from the loop. I'm guessing a microscopic leak that evaporated as fast as it leaked.
- Two power supplies -- One randomly exploded. Just started making popping noises and shooting sparks out the back. At autopsy, I determined that the fan failed (it was hard to spin manually) and it overheated. The other was just being overloaded. I had just gotten a new GPU and the PSU wasn't big enough for it.
- Several case fans.
Never seen a CPU or motherboard failure, even when my AIO water cooler was failing and my CPU was constantly at thermal limits. Never had an SSD/NVMe failure.
After the second, HP replaced all that we had from that batch.
HDDs fail every 3-6 months, there's a lot of them (3-4PB).
One PSU.
One CPU/motherboard. (I let the HP engineer handle that one.)
One NVMe drive.
I think over the decades I got one RAM disk fail. And another one didn't fail but, although I bought a pack of four sticks (2x x2), was detected (by memtest, when building the rig), as having another model name than the three others: even the shop who sold me the RAM was confused (don't know how that happened).
The usual component that I've seen fail the most would be the PSU? (before I started buying quality ones). One NVMe drive (an "ADATA") died on me even though it was nearly new.
PC assembled with quality parts usually last a very long time.
The only thing i've never seen fail is the Case lol. Air cooler too if you don't count the fan.
Common scenario:
200 PCs at least 6 years old, often over a decade old, come in on pallets.
We pick up the first 50, stack them on benches, and power them on.
At least 10 wont post.
Of that 10, reseating memory and other little jobs fixes 5.
Of the remaining 5, replacing or removing memory gets 4/5 them to post.
Testing those sticks in other PCs confirms the RAM is dead.
Into the scrap it goes.
The remaining PC is dead mobo, PSU or needs a new CMOS battery.
...granted though I've also seen a backpane on a server fail. I didn't think that was even possible until that point.