Intel is selling defective 13-14th Gen CPUs
alderongames.com
alderongames.com
If it’s as they say, switching to AMD shouldn’t lead to 100x fewer crashes, it should lead to no crashes.
Lumping both 13th and 14th gen hardware together makes it harder to take the claims as easily at face value. 13th gen have been around long enough that you’d think they’d be able to dig up a kernel mailing list bug report or some other corroboration of their accusation.
I clicked expecting basically an undocumented errata with specific, 100% reproducible instructions and a deep dive into the internal architecture changes that could have caused this. I’m tempted to flag the submission, honestly. Claims of such magnitude demand at least some baseline evidence.
No, this does not follow, at all. Few things in life are monocausal and systems (here the aggregate of hardware and software) tend to have more than one bug. If one bug dominates and you remove it then of course you're still left with residuals.
It doesn't have to be intel's fault though. It could be mobo vendors defaulting to unsafe settings.
Sure. However workstation and server class motherboards have a much lower risk of not following Intel recommendations on voltages and clock speeds.
The error rate on the workstation boards are pretty high, on the order of a failure per week for 50% of workstations they were collecting telemetry from.
I fully agree that demanding Intel recall all of the CPUs at this point is premature. They've done nowhere near enough research to pin the problem on the processor and not their code.
Well, ... they are. This has been being discussed for a while. It came up a bit ago, many people were seeing it in games with a weird error about low VRAM which apparently wound up being linked to Intel's 13th and 14th gen processors. Exactly why is anyone's guess.
https://www.tomshardware.com/pc-components/cpus/nvidia-blame...
Level1Techs is relatively well-trusted, and they seem to have some contacts that can also corroborate these issues.
500 watt power supply becomes less efficient over time, now it's not quite 500 watts, parts of the chip brown out. Undefined behavior.
I also don't know exactly what failure mode you are suggesting with this (voltage rail out of spec?) but I think unless it starts feeding too much voltage into the CPU it is probably unlikely to be able to cause damage. (Instability, sure.)
It also doesn't explain why Intel's 12th generation CPUs are still running strong, along with most of AMD's lineup, by comparison. Don't get me wrong, every SKU of CPU has its issues, but all of the major issues with other CPU lines we have explanations for; the issues with Intel Raptor Lake are thusfar unexplained. Intel has blamed motherboard manufacturers quite extensively, but so far the evidence of this is pretty uncompelling, because the failures can apparently also be seen in abnormally high numbers on motherboards that never had unreasonable configurations.
The PSU (notably in this case the M/B VRM, not the ATX PSU) can't really trigger brownouts. 1.05V is 1.05V, and if asked to do so, the PSU will deliver 1.05V. It will just get hotter while doing so over time (and specifically: capacitors degrading.) The voltage references are quite stable over time, and so are the window comparators that ensure "1.05V" is between, say, 1.035V and 1.065V. (Voltage references have essentially continuous engineering history back to the 50s≈60s, it's a "solved" problem.)
Even with generously specced power supplies, server class hardware, redundant power supplies, and enough smarts to stage the drive spin ups to avoid a large peak inrush current when 16 drives try to spin up at once.
Now I'm thinking of these HPE over-over-over-specced PSUs I run and y'all got me worried.
Sure power supplies provide accurate power ... till they can't. Unlike peak power that starts to degrade from new, the voltages will hold as long as the PS can manage it.
And again, the ATX PSU doesn't really matter here, the M/B VRM does. Even if the ATX PSU has a failure mode where it delivers 10V instead of 12V — you can still make 1.05V out of 10V. Decoupling is spread all the way too, so short peaks don't explain a lot either.
Lastly, it's of course possible the VRM is misdesigned. But this issue is seen across mainboards from at least 3 vendors… what is the likelihood of all of them getting it wrong in a very similar way? The primary way for that to happen is if they are all implementing Intel's reference design, in which case it's Intel's fault again anyway and we've looped back to the origin of the discussion…
https://www.youtube.com/watch?v=gTeubeCIwRw
similarly, the oodle maintainers only found it to affect a small % of machines (but reliably, of course).
https://news.ycombinator.com/item?id=39486930
The alderon_matt thing is actually carefully worded to say affected machines invariably degrade... but they don't say how many machines are affected. It's deceptively worded, intended to be read the way you read it, but actually not saying that 100% of machines are affected either.
at the end of the day game developers are just people too, and they're equally frustrated with the process/not willing to cut intel slack. but like, this is shades of the Ryzenfall thing (which, remember, contrary to early takes did have several real severity 9.0 issues behind it, including a PSP jailbreak and a UEFI module signing bypass, as well as memory encryption bypass and some other goodies - AMD doesn't issue AGESA patches for "root password lets you do root things"). People have moved past the facts and into the "trying to generate noise to get attention" on levels that are not supported by the facts.
https://www.youtube.com/watch?v=QuqefIZrRWc
the Alderon Games assertion, specifically, is very aggressively/deceptively worded and is not supported by the broader facts, nor does he have any real idea what can or can't be patched. as wendell notes - all kinds of "unfixable" hardware bugs are indeed fixed in microcode, routinely. that's why it's there, because of the pentium FOOF bug.
this isn't to say there isn't a degradation problem, there clearly is something (or more than one something) going on. but the alderon games press release is very over-the-top compared to all the rest of the evidence. GN is pointing at 10-25% and this is very consistent across a number of different operators, boards, environments, etc. That's the number Dell is getting in their validation (depending on the department/how stringent they validate), that's the number Wendell is seeing in operation in datacenter error reports. And they're thinking possibly some kind of fab defect affecting those units, basically (oxidation of the vias).
What's making news is that Intel is starting (in at least one case) to deny RMA replacements after the first, despite shipping defective CPUs as replacements.
Yes: games are notorious for having poor multithreading and hitting one core with very high load. And unlike servers where you're rarely the only workload, the rest of the CPU will be rather idle on a desktop/gaming system. This pushes the CPU to a much more "imbalanced" mode (putting all the power and heat into a very small area) than is common elsewhere.
No?
There’s always multiple sources of crashes. Switching from Intel may eliminate 100% of one crash source, but not 100% of all crashes.
One of my favorite stories is that ArenaNet once got Guild Wars servers so stable that any crash could be reliably attributed to hardware failure. Usually memory IIRC.
(NB: random hardware failures like we're talking about here. CPU design/logic erratas can and do look like code bugs; but that's not what's being discussed here.)
I totally agree, and this entire HN post is about the fact that these Intel CPUs probably have some yet-undiscovered erratas. Best theory at this point is that it's power/thermal related though.
> there will be a large number of combinations of instructions/actions that should be avoided
As a matter of fact, reading those erratas is part of my job, and no, it is very rare for such erratas to have the effect you're implying. Rare enough that this issue that is seen here would already have been matched against an errata.
I don't see anything in Intel's current errata — other than their in-progress responses to these reports — that could explain this. Here's the errata sheet for 13th/14th gen: https://edc.intel.com/content/www/us/en/design/products/plat... … do you see anything?
That also doesn't exist on x86 since it is a TSO platform.
The other thing is that even in the extremely unlikely case that some new behavior in 13th/14th gen Intel CPUs is exposing a bug in either code or compiler that has just been lurking there for who knows long, and doesn't affect any other CPU at all…
… with the way the world works, it'd still be a CPU bug, even if there was a document somewhere that says the CPU is right. x86 is shoved down the compatibility road so far, it just doesn't matter what some paper says if you break compat. (For user space at least, kernel is not quite there.)
Edit: 14th gen top end parts are basically 13th gen parts with different binning and higher clocks.
This is pretty damn reasonable actually. The 14th gen Intel desktop hardware is for all intents and purposes, essentially the same as the 13th gen hardware. It behaves nearly identically, it benchmarks nearly identically.
[1] https://www.epicgames.com/help/en-US/c-Category_Fortnite/c-F...
He's reached out to folk who gather game telemetry and got some interesting data to play with.
Based on Fortnite’s suggested fix in their help, it sounds like there there is some flaw in the Intel hardware that makes it crash later in life when using the shipping voltage, which is maybe its most important parameter. When that is forced to be higher the part gets much hotter and performance per watt declines.
The argument is that Intel is selling chips that at their out-of-the-box recommended use patterns can wear out extra quickly, and by default are over-volted. So it's not a "flaw" so much as not having a great engineering safety margin for premature wear.
There are at least 2 causes and failures modes, intel has already confirmed eTVB=off was a problem and could cause degradation due to excessive heat. The other suspect right now is the ring bus degrading, perhaps.
It's just laziness. The RAD game tools statement/article[1] has much more detail. I'm not sure to what degree the people behind the current article verified that they're seeing the same issue, but there is definitely a issue (apparently power/scaling/temperature related).
I’m done with Intel.
First of all the 13th and 14th gen are nearly the same CPUs. Slight differences in clock, TDPs, and core counts, but all using the identical raptor lake cores. Intel's been widely ridiculed for calling them 14th gen when there's no difference in the cores.
There have been various posts about these stability problems going back months on numerous forums. Often things like BSOD, not enough VRAM errors, decompression errors, failure to load game assets, which looks like a NVMe error, etc.
Part of the reason it's flying under the radar is a combination of factors. It's not all chips, it's not a specific instruction, and the behavior changes over time by getting worse. Sadly if a game or OS crashes nobody is terribly surprised, which can make hardware problems less obvious.
From what I can tell there Intel clock/voltage/temp curves were pushed, trying to be competitive with AMD despite a disadvantage Intel has with their 10nm fabs. Not only does this cause errors, but it causes damage, so the error rate increases over time.
The best source found seems to be the telemetry built into unreal engine that provides an unbiased report on crashes (but not hangs, BSoD aren't reported). Sure some gamers overclock, have poor airflow, poorly applied headsink goo, etc. But I found the following pretty compelling since it's based on workstation hardware, without overclocking, using the W series workstation chipset not the Z series consumer chipset which allows overclocking:
In a test population of more than 210 W680-based systems, 47.1% of these systems experience at least one incident of instability over a 168 hour test window. This distribution is the same to within 0.4% between Asus brand W680 and Supermicro W680 based boards.
For people running these boards on the server side they actually are charging more because of support issues related to CPU replacements.Certainly a failure of 50% per week per node is crazy high and unacceptable on workstation class hardware.
>If it’s as they say, switching to AMD shouldn’t lead to 100x fewer crashes, it should lead to no crashes.
As many others have pointed out, this is baloney. Computers are not perfect machines and even if they were, gathering aggregate user data is also imperfect.
> I clicked expecting basically an undocumented errata with specific, 100% reproducible instructions and a deep dive into the internal architecture changes that could have caused this.
I don't know why you would expect that. You know what an intermittent bug is, this feels like disingenuous reasoning. Furthermore, there are many other people who are noticing this, documenting it, and attempting to report on it.
Intel has a Pretty Big Problem (by Level1Techs): https://www.youtube.com/watch?v=QzHcrbT5D_Y
Intel's CPUs Are Failing, ft. Wendell of Level1 Techs (by Gamers Nexus): https://www.youtube.com/watch?v=oAE4NWoyMZk
It is a bit odd that after months of this being well-known problem Intel still hasn't released a root cause. Is it an instruction sequence? Is their code subtly wrong in a way that works sometimes on these chips but always on other chips? Are some processors defective?
There may also be some AI chip design shenanigans going on as well, which would further exacerbate the human comprehension difficulty scale.
That's a really high bar for a non-Intel investigation. The Pentium III 1.13GHz issue 24 years ago had 100% reproduceable instruction, but even that didn't include any deep dive into the internal architecture because that just isn't information that people have.
https://www.intel.com/content/www/us/en/support/articles/000...
So you kind of have to go all in on Turbo boost 3.0 as the issue and then explain why the non-K variants of the i7 and i9 don't seem to have the same frequency of occurrence.
He was able to identify a large group of CPUs used in W680 chipset motherboards (i.e. using stock Intel power curves, never seeing a lick of overclocking/overvolting) which exhibited issues.
It would seem enthusiast ricing is, at best, an aggravating factor to some other underlying cause.
If the latter, I am afraid I don't see the point you are hinting at and need more clarification.
What you are getting at is Turbo boost, which is not a new thing. We can split hairs over the incremental enhancements to Turbo boost over the years of course.
Nvidia's been flat out telling people complaining about games crashing on 13th and 14th gen CPUs to 'talk to Intel':
https://www.techspot.com/news/102611-nvidia-advises-crash-pr...
If you overlooked Ryzen because of issues like memory latency in the early days, I encourage looking at them again. If you do your homework, you can find motherboards that support ECC ram to varying degrees (ie from "it'll run but with ECC disabled" to "error correcting but doesn't report it" to "full support, just not listed as such".) I believe AM4-socket ECC support is much more prevalent, if that's important to you.
Familiarize yourself with AMD's Precision Boost Overdrive, which can be used for overclocking and/or undervolting. Unlike AMD's graphics cards (which are space heaters, though the current 7xxx series is an improvement), Ryzen processors are fairly efficient in stock form; more so if you do even a bit of mild undervolting.
Don't waste your time with the stock coolers supplied with any of the chips. They'll keep the chip cool enough, but are noisy as hell compared to a well-performing $30 dual-fan air cooler (GamersNexus has found several in their reviews.)
Keep your BIOS up to date for AGESA (AMD microcode) improvements, and keep your AMD chipset drivers up to date as well.
I don't think I should have issues?
I've ran a 13900K from launch, no issues.
Given what little insight they have now, the crash-triggering fault could just as likely be in the peculiarities of their own code, some flaw in their build toolchain, some flaw in a runtime library, etc
With all respect to the folks at this studio, this statement reads like a junior dev throwing up their arms and blaming others over a bug they can't wrap their head around.
Even if they have a hunch that it's a CPU bug and can't afford to prove it, a more professional approach would be to perform the hardware change to AMD and toss out a more curious and less authoritative statement -- "Anybody else seeing more crashes on these processors? Have you figured out what from?" vs "These are broken and should be recalled".
Idiot: There's something wrong with this CPU
Midwit: Nooo it's never the CPU if you think it's the CPU you must be a junior dev
Genius: There's something wrong with this CPU
And there could certainly be a CPU issue here. My critique is mostly about them trying to make such a strong statement about it with little specific evidence, despite not needing to make a statement at all.