I fully agree that demanding Intel recall all of the CPUs at this point is premature. They've done nowhere near enough research to pin the problem on the processor and not their code.
I fully agree that demanding Intel recall all of the CPUs at this point is premature. They've done nowhere near enough research to pin the problem on the processor and not their code.
Well, ... they are. This has been being discussed for a while. It came up a bit ago, many people were seeing it in games with a weird error about low VRAM which apparently wound up being linked to Intel's 13th and 14th gen processors. Exactly why is anyone's guess.
https://www.tomshardware.com/pc-components/cpus/nvidia-blame...
Level1Techs is relatively well-trusted, and they seem to have some contacts that can also corroborate these issues.
500 watt power supply becomes less efficient over time, now it's not quite 500 watts, parts of the chip brown out. Undefined behavior.
I also don't know exactly what failure mode you are suggesting with this (voltage rail out of spec?) but I think unless it starts feeding too much voltage into the CPU it is probably unlikely to be able to cause damage. (Instability, sure.)
It also doesn't explain why Intel's 12th generation CPUs are still running strong, along with most of AMD's lineup, by comparison. Don't get me wrong, every SKU of CPU has its issues, but all of the major issues with other CPU lines we have explanations for; the issues with Intel Raptor Lake are thusfar unexplained. Intel has blamed motherboard manufacturers quite extensively, but so far the evidence of this is pretty uncompelling, because the failures can apparently also be seen in abnormally high numbers on motherboards that never had unreasonable configurations.
The PSU (notably in this case the M/B VRM, not the ATX PSU) can't really trigger brownouts. 1.05V is 1.05V, and if asked to do so, the PSU will deliver 1.05V. It will just get hotter while doing so over time (and specifically: capacitors degrading.) The voltage references are quite stable over time, and so are the window comparators that ensure "1.05V" is between, say, 1.035V and 1.065V. (Voltage references have essentially continuous engineering history back to the 50s≈60s, it's a "solved" problem.)
Even with generously specced power supplies, server class hardware, redundant power supplies, and enough smarts to stage the drive spin ups to avoid a large peak inrush current when 16 drives try to spin up at once.
Now I'm thinking of these HPE over-over-over-specced PSUs I run and y'all got me worried.
Sure power supplies provide accurate power ... till they can't. Unlike peak power that starts to degrade from new, the voltages will hold as long as the PS can manage it.
And again, the ATX PSU doesn't really matter here, the M/B VRM does. Even if the ATX PSU has a failure mode where it delivers 10V instead of 12V — you can still make 1.05V out of 10V. Decoupling is spread all the way too, so short peaks don't explain a lot either.
Lastly, it's of course possible the VRM is misdesigned. But this issue is seen across mainboards from at least 3 vendors… what is the likelihood of all of them getting it wrong in a very similar way? The primary way for that to happen is if they are all implementing Intel's reference design, in which case it's Intel's fault again anyway and we've looped back to the origin of the discussion…
https://www.youtube.com/watch?v=gTeubeCIwRw
similarly, the oodle maintainers only found it to affect a small % of machines (but reliably, of course).
https://news.ycombinator.com/item?id=39486930
The alderon_matt thing is actually carefully worded to say affected machines invariably degrade... but they don't say how many machines are affected. It's deceptively worded, intended to be read the way you read it, but actually not saying that 100% of machines are affected either.
at the end of the day game developers are just people too, and they're equally frustrated with the process/not willing to cut intel slack. but like, this is shades of the Ryzenfall thing (which, remember, contrary to early takes did have several real severity 9.0 issues behind it, including a PSP jailbreak and a UEFI module signing bypass, as well as memory encryption bypass and some other goodies - AMD doesn't issue AGESA patches for "root password lets you do root things"). People have moved past the facts and into the "trying to generate noise to get attention" on levels that are not supported by the facts.
https://www.youtube.com/watch?v=QuqefIZrRWc
the Alderon Games assertion, specifically, is very aggressively/deceptively worded and is not supported by the broader facts, nor does he have any real idea what can or can't be patched. as wendell notes - all kinds of "unfixable" hardware bugs are indeed fixed in microcode, routinely. that's why it's there, because of the pentium FOOF bug.
this isn't to say there isn't a degradation problem, there clearly is something (or more than one something) going on. but the alderon games press release is very over-the-top compared to all the rest of the evidence. GN is pointing at 10-25% and this is very consistent across a number of different operators, boards, environments, etc. That's the number Dell is getting in their validation (depending on the department/how stringent they validate), that's the number Wendell is seeing in operation in datacenter error reports. And they're thinking possibly some kind of fab defect affecting those units, basically (oxidation of the vias).
What's making news is that Intel is starting (in at least one case) to deny RMA replacements after the first, despite shipping defective CPUs as replacements.
Yes: games are notorious for having poor multithreading and hitting one core with very high load. And unlike servers where you're rarely the only workload, the rest of the CPU will be rather idle on a desktop/gaming system. This pushes the CPU to a much more "imbalanced" mode (putting all the power and heat into a very small area) than is common elsewhere.