> Lumping both 13th and 14th gen hardware together makes it harder to take the claims as easily at face value. 13th gen have been around long enough that you’d think they’d be able to dig up a kernel mailing list bug report or some other corroboration of their accusation.
First of all the 13th and 14th gen are nearly the same CPUs. Slight differences in clock, TDPs, and core counts, but all using the identical raptor lake cores. Intel's been widely ridiculed for calling them 14th gen when there's no difference in the cores.
There have been various posts about these stability problems going back months on numerous forums. Often things like BSOD, not enough VRAM errors, decompression errors, failure to load game assets, which looks like a NVMe error, etc.
Part of the reason it's flying under the radar is a combination of factors. It's not all chips, it's not a specific instruction, and the behavior changes over time by getting worse. Sadly if a game or OS crashes nobody is terribly surprised, which can make hardware problems less obvious.
From what I can tell there Intel clock/voltage/temp curves were pushed, trying to be competitive with AMD despite a disadvantage Intel has with their 10nm fabs. Not only does this cause errors, but it causes damage, so the error rate increases over time.
The best source found seems to be the telemetry built into unreal engine that provides an unbiased report on crashes (but not hangs, BSoD aren't reported). Sure some gamers overclock, have poor airflow, poorly applied headsink goo, etc. But I found the following pretty compelling since it's based on workstation hardware, without overclocking, using the W series workstation chipset not the Z series consumer chipset which allows overclocking:
In a test population of more than 210 W680-based systems, 47.1% of these systems experience at least one incident of instability over a 168 hour test window. This distribution is the same to within 0.4% between Asus brand W680 and Supermicro W680 based boards.
For people running these boards on the server side they actually are charging more because of support issues related to CPU replacements.
Certainly a failure of 50% per week per node is crazy high and unacceptable on workstation class hardware.