Intel Processor Instability Causing Oodle Decompression Failures
radgametools.com
radgametools.com
https://forum.level1techs.com/t/amd-threadripper-3970x-under...
HN discussion: https://news.ycombinator.com/item?id=22382946
Ended up investigating the issue with AMD for several months, was generously compensated by AMD for all the troubles (sending motherboards and CPUs back and forth, a real PITA), but the outcome is that I've been running since then with a custom BIOS image provided by AMD. I think at the end the fault was on Gigabyte's side.
They sent us custom BIOSes until it got stabilized and said they'll put the patch in the following BIOS releases.
The thing is neither Intel nor AMD nor Supermicro can test edge cases at max usage in niche environments without paying money, but they would really love to claim with backup they can be integrated for such solutions. If Intel wants to test stuff in space for free they have to cooperate with NASA; the alternative is in-house launch.
If users pay seven figures+ it might make sense.
I was doing reinforcement learning on that system and it was always crashing, I spent quite a bit of time trying to find the problem, swapped the CPU for a 13700kf I was using in another PC, the problem was solved.
So I contact Intel to start the RMA process, Intel said that the MSI motherboard I was using doesn't support Linux, I emailed them the official Intel GitHub repo with the microcode that enables the support, they switched agents at that point but I was clear to me at that moment that Intel was trying their best to avoid the RMA, luckily I live in Europe, so I contacted my local consumer protection agency and did the RMA through them, in the meanwhile I saw a good offer for a 7950x + motherboard in an online retailer, bought it and sold in the second market my old motherboard and the RMA 13900k when I got it.
Not buying Intel ever again, I was using Intel because they sponsor some projects in DS but damn.
The hyperthread and c-state stuff, eh, if you want to run code that might be a virus you will have to limit your system. I dunno. It would be a shame if we lost the ability to ignore that advice. Most desktops are single-user after all.
these chips that have been specially binned because they are supposedly stable at those frequencies (within an envelope set by intel)
if intel can't get it to work they shouldn't be selling these chips at all
People want to overclock. Gamers want to see big numbers. If gamers don't do it their motherboard vendors will. It's not a market over which Intel is going to have much control, really.
Note that you don't, in general, see this kind of silly edgelord clocking in the laptop segments.
Out of the box default overclocking is not, this aspect should be policed.
This issue doesn't affect every such machine, but both people affected by the issue that consented to run tests for us still had the issue reproduce after flashing BIOS to current and with BIOS default settings for absolutely everything.
Among the settings enabled by default on some boards: current limit set to 511 amps (...wat), long duration power limit set to 350W (Intel spec: 125W), short duration power limit also set to 350W (Intel spec: 253W), "MultiCore Enhancement" which is extra clock boosting past what the CPUs do themselves set to "Auto" not "Off", and some others.
So, you are trusting all web pages you view? Because these are unknown code running on your box which probably has some beefy private data.
I imagine people doing e.g. heavy number crunching might want something similar.
I'm not affiliated with NoScript. I just think it's insane that we run oodles of code to display web pages.
I thought the entire premise of overclocking was that it's not officially supported and it may break things.
The whole point is that you're not paying for it and it's entirely at-risk.
Because if you do want a higher level of guaranteed performance, you do need to pay for a faster chip (if it exists).
It’s fair for the end user who bought a motherboard that promises a higher clock speed to expect that clock speed.
If you can provide links, I'd be curious to see what guarantees they make. "What's fair" depends very specifically on what language they use.
They have now entered the AI bubble with
https://www.asus.com/microsite/motherboard/Intelligent-mothe...
MSI has a similar setting, although I don't know exactly what models have it nor what it's called
Tell that to anyone who paid extra for a K-series Intel chip.
Its cheaper to have a single production line and then lock off features.
As crazy as i sounds it actually cost a little bit more to produce inferior version sold at cheaper prices.
The overclocking was a 'premium' feature due to possibility of melting the chip. But nowadays the temp sensors cut power to prevent catastrophic failure.
Also worth mention the downside to upclocking voltage is increased physical degradation of cores, ie lower lifespawn of cpu.
These chips require motherboards to function, and these unlocked chips get their configuration from the motherboard. There’s no analogous entity to Ferrari the company here, it is like you bought an engine from one company, a gearbox from another, and the gearbox had a “responsiveness enhancement” setting that always redlined your RPMs or something (I don’t know cars).
I don't know what you want Intel to do here. They tell you upfront what the power and clock limits are on the parts. But the market has a three decade history of people pushing the chips a little past their limit for fun and profit, so they "allow" it even if they know it won't work for everything.
- If I buy parts that are certified to work together, and I use them according to their respective manuals, they should work as specified.
- If I desire to manually change or customise something, I should be able to modify whatever I'd like.
- As soon as my changes go outside the certified range, I'm liable myself. But as long as I'm within of the certified range, warranty should still apply and the product should continue working as specified.
That's a blanket statement, and wrong. Ferrari doesn't allow unlicensed modifications of their cars.
You are able to customize many vehicles to your liking. And just like you can choose the options before sale, you're free to replace one official part with another official part after sale as well.
From what I can tell, this is limited to rims, tires, brakes, seats, passenger display, and other similar configuration options, though.
Search for "3000hp lambo" on youtube and you'll see what modification actually means.
We're talking about using one intel-certified part with another intel-certified part using intel-certified default settings.
I don't have a problem with end users experiencing instability once they manually overclock (that's how it goes), but CPU + mainboard combinations experiencing typical OC symptoms with out-of-the-box settings is just not OK.
This appears to be an arms race between mainboard vendors all going further and further past spec by default because it gives better benchmark and review scores and their competition does it. Intel for their part are themselves also dialing in their parts more aggressively (and, presumably although I don't know for sure, with smaller margins) over time, and they are for sure aware that this is happening, because a) even had they not known already (which they did) they would have learned about this months ago when we first contacted them about this issue, b) technically out of spec or not, as long as it seems to work fine for users and makes their parts look better in reviews, they're not going to complain.
However, it turns out, it does not work fine for at least some small fraction of machines. I have no idea what that percentage is, but it's high enough that googling for say "Intel 13900K crash" yields plenty of relevant results. Some of this will be actual intentional overclockers but, given how boards default to some extend of out-of-spec overclocking enabled, it's unlikely to be all of them.
Meanwhile we (and other SW vendors) are getting a noticeable uptick in crash reports on, specifically, recent K-series Intel CPUs, and it's not something we can sanely work around because the issue manifests as code randomly misbehaving and it's not even when doing anything fancy. The Oodle issue in particular is during LZ77-family decompression, which is to say, all integer arithmetic (not even multiplies, just adds, shifts and logic ops), loads/stores and branches. This is the bare essentials. If it was an issue with say AVX2, we could avoid AVX2 code paths on that family of machines (and preferably figure out what exactly is going wrong so we can come up with a more targeted workaround). But there is no sane plan B for "integer ALU ops, load/stores and branches don't work reliably under load". If we can't rely on that working, there is not enough left for us to work around bugs with!
I realize this all looks like finger-pointing, but this is truly beyond our capacity to work around in a sane way in SW, with what we know so far anyway. Maybe there is a much more specific trigger involved that we could avoid, but if so, we haven't found it yet.
Either way, when it's easy to find end user machines that are crashing at stock settings, things have gone too far and Intel needs to sit down with their HW partners and get everyone (themselves included) to de-escalate.
You think that might have something to do with you having put "Intel Processor Instability" in the title of a whitepaper on an issue that you already root caused to motherboard settings? I mean, did you want to troll a big flame war? Because this is how you troll a big flame war.
I don't think it's unreasonable to call that Intel's problem, maybe not in terms of culpability (but truly, nobody cares) but definitely in the sense this is doing damage to their brand. If the mainboards are all out of spec then they need to talk about this publicly, rein them in, start a certification program, whatever. Being publicly completely fine with this as long as it results in good review scores but then starting to go "well actually..." when there's stability issues on a small fraction of sold units is not a good look.
You didn't call it Intel's problem. You said Intel CPUs were "unstable", which simply isn't true. If your title was "Intel doesn't police default BIOS clocking", we wouldn't be this far down in the senseless thread about semantics. (Though to be fair, you wouldn't have been on the front page as long either, so maybe that's as intended.)
Out of the box with default settings, it was pushing 320W through the CPU in stress tests.
I use my machine for FPGA compiles so I need reliability. I learned that ASUS Multicore Enhancement is not the only thing that must be disabled, you must manually enter the power limits.
Now my compiles take exactly the same length of time but use at least 100W less power.
I am glad to know that with your field data, I've inadvertently sidestepped a potentially catastrophic bug. I don't want to release an FPGA bitstream to users with flipped bits. And the FPGA tools already crash on their own enough.
The problem is that Intel has normalized it so much that all their high end CPUs do this, and apparently do it often. It's not unexpected that they might be too close to the point where things are melting, so to speak.
I'd rather slower and more stable any day - I chose a Ryzen 7900 over a 7900X intentionally - but that isn't what all the marketing out there is trying to sell. The fancy motherboards, the water coolers, the highly clocked memory all account for lots of markup, so that's what's marketed. I'm not a fan.
It is worth noting a distinction between the terms "overclocking" and "turbo clocking". "Overclocking" has traditionally meant running the clock "over" the rating. "Turbo clocking" is now built in to almost every CPU out there. One technically can void your warranty, whereas the other doesn't.
Since we're mostly technical people here, we should use the appropriate term where the context makes that choice more accurate. It's like virus and Trojan - we SHOULD be technically correct, but that doesn't mean highly technical people aren't still calling Trojans viruses now and then.
This "I can run a core at a faster speed" is a documented feature so not really overclocking.
As an example of motherboard manufacturers going outside specifications, my MSI motherboard has a built-in option to change BCLK, which is the clock reference for the entire PCIe bus. Changing it not only overclocks the CPU, but also the GPU's connection (not the GPU itself), as well as the NVMe SSD.
This was so not-endorsed by Intel that they quickly pushed microcode that shuts the CPU down if it detects BCLK tampering.
In response, MSI added a dropdown that allows you to downgrade the microcode of the CPU.
So yeah. Very not within specifications.
If I run the chip in a way not documented by the manufacturer, or modify the ECU to allow the turbo to generate more boost, those are both unsupported modifications, and I'd consider either of those "overclocking"
I wonder if the old style “bump it up and memtest” type overclocking would catch this. Actually, what is the good testing tool nowadays? Does memtest check AVX frequencies?
I don't know how common this is across the whole population of PC buyers, but personally, I have for sure bought K-series parts then not clocked them past their stock settings, trusting that they are rated for it and deeply uninterested in any OCing past that. (I prefer my machines stable, thank you very much.)
He was so happy with the speed. Would not stop telling everyone, and talking about it. Yet as I watched him demo it, it rebooted every minute or so. Most unstable thing ever.
Sure it booted in 2 seconds, and he just went about his merry way, but.. what?! Guy could have still overclocked a little less and had stability, but nope.
Some overclockers are weird.
https://news.ycombinator.com/item?id=39479081
Somehow people think that it's a strawman, but people like parent comment actually think and post like this lol
It was a nightmare to get running stable. None is the default settings the motherboard used worked. Games crashed, kernel and emacs compiles failed.
End result I had to cap turbo to 5.4ghz on a 6ghz chip, and enable settings that capped max watts and temperature for throttling to 90c.
System seems stable now. Can get sustained 5.4ghz without throttling and enjoying games at 120fps with 4k resolution.
Even though it is working I do feel a way about not being able to run the system at any of the advertised numbers I paid for.
You know how ISPs used to sell "up to X Mbps"? Same idea. Your chip will turbo boost "up to 6.00 GHz".
It's basically automated overclocking, and as you learned, sometimes it can't even do it in a stable fashion. Some of those chips will never clock "up to 6.00 GHz" but they didn't lie. "up to"
And laptops have another layer of bullshit, because the theoretical boost clocks the chip is capable of will in practice be limited by the power delivery and cooling provided by that specific machine, and the OEMs never tell you what those limits are. So they'll happily take an extra $200 for another 100MHz that you'll never see for more than a few milliseconds while a different model with a slower-on-paper CPU with better cooling can easily be more than 20% faster.
Its hard to cool these new chips. AMD included.
So unless you have powerful cooling you will hit the thermal limit.
enable settings that capped max watts and temperature for throttling to 90c.
You were going above 90C before???My first thought is that seems insane, but is apparently normal for that chip, according to Intel: "your processor supports up to 100°C and any temperatures below it are normal and expected"
https://community.intel.com/t5/Processors/i9-14900K-temperat...
That is just wild though. On one hand you should obviously get the performance that was advertised and that you paid for. On the other hand IMO operating a CPU at 90-100C is just insane. It really feels like utter desperation on Intel's part.
I would be curious what kind of cooling setup you have.
No it isn't, the manufacturer literally says it's normal! I think people who spend as much money on cooling setups as the chip are the insane ones.
My favorite story: I once put Linux on a big machine that had been running windows, and discovered dmesg was full of thermal throttling alerts. Turns out, the heatsink was not in contact with the CPU die because it had a nub that needed to occupy the same space as a little capacitor.
I'd been using that machine to play X-plane for over two years, and I never noticed. It was not meaningfully slower: the throttling events would only happen every ten or so seconds. I'm still using it today, although with the heatsink fixed :)
I have a garage machine with a ca. 2014 Haswell that's been running full tilt at 90C+ for a good bit of its life. It just won't die.
I have a garage machine with a ca. 2014 Haswell
that's been running full tilt at 90C+ for a good
bit of its life. It just won't die.
First of all: that's awesome.Second: what's it doing spending that much time "full-tilt?" Just curious. Sounds like something interesting.
Right now, I use it as a 200W space heater I can turn off and on via SSH, just running busyloops. It works well enough for a single car garage in California.
First, the temperature sensors got a lot better. Previously you only had one sensor per core/cpu, and it was placed wherever there was space - nowadays it'll have dozens of sensors placed in the most likely hotspots. A decade ago a 70C temp meant that some parts of the CPU were closer to 90C, whereas nowadays a 90C temp means the hottest part is actually 90C.
Second, the better sensors allow more accurate tuning. While 100C might be totally fine, 120C is probably already going to cause serious damage. The problem here is that you can't just rely on a somewhat-distant sensor to always be a constant 20C below the peak value: it's also going to be lagging a bit. It might take a tenth of a second for that temp spike in the hotspot to reach the sensor, but in the time between the spike starting and the temp at the sensor raising enough to trigger a downthrottle you could've already caused serious damage. A decade ago that meant leaving some margin for safety, but these days they can just keep going right up to the limit.
It's also why overclocking simply isn't really a "thing" anymore. Previous CPUs had plenty of safety margin left for risk-takers to exploit, modern CPUs use up all that margin by automatically overclocking until it hits either a temperature limit or a power draw limit.
Similarly, processors convert thermal headroom to performance, until they run out of thermal headroom. So if you improve the cooling on a processor that has work to do (rather than sleeping), it will use up that cooling and perform better, rather than performing the same and running cooler.
(Mobile processors operate differently, since they need to maintain a much tighter thermal envelope to not burn the user. And processors can also target power levels rather than thermals. But when a processor is in its default "run as fast as possible" mode, its normal operating temperature will be close to the 100C limit.)
This is really no excuse to have Windows eating 12GB of RAM with an Edge (with 2 static tabs), excel some antivirus and Teams running. This (12GB) is just sick.
If you have free ram the OS will cache all the recently accessed things. It costs effectively zero to cache them when you access them, and it costs effectively zero to evict from the cache if you actually need the memory for "useful" work. And if the cached thing is useful in the mean time, you win because RAM is still orders of magnitude faster than SSD. None of the scenarios are a perf loss. The only scenarios are "win" or "no gain."
"But won't the next application I load, load faster if there is free RAM?"
If you are asking this question, please refer back to the previous concept -- "it costs effectively zero to evict from the cache if you actually need the memory for 'useful' work." This is where people typically get confused. They think it takes work to evict things from the cache, just like it would require work to remove frequently used objects from your physical desk if you needed to free up space.
Once you understand that the rest should fall into place.
Unless you take reliability into account.
It might look insane compared to what you were used
to seeing a decade ago, but it's actually not that crazy.
I'm comparing it to a mid-range system I built last year: an Intel i5-13400F with a $60 closed-loop cooler.It doesn't go past 45C even during Prime95 runs using all cores. And the 120mm fans aren't exactly howling either.
First, the temperature sensors got a lot better. Previously
you only had one sensor per core/cpu, and it was placed wherever
there was space - nowadays it'll have dozens of sensors placed
in the most likely hotspots.
That's awesome to know. Thanks for the explanation.How much margin is really there?
Therefore, 1.1x the frequency at the high end (where switching power dominates) is 1.33x the power draw.
Those final few hundred MHz really hurt. Conversely, that's also why you see "Eco" power profiles with a major reduction in power draw that cost you maybe 5-10% of your peak performance.
I'm curious about what you are using for cooling, as 90C at 5.4ghz seems way off compared to what I am seeing on my processor, but it could just be that I'm not pushing my processor quite as hard even with the higher clock rate.
https://www.tomshardware.com/reviews/intel-admits-problems-p...
One decision that Zstd made was to include only a checksum of the original data. This is sufficient to ensure data integrity. But, it makes it harder to rule out the decompressor as the source of the corruption, because you can't determine if the compressed data is corrupt.
https://www.radgametools.com/granny.html https://www.radgametools.com/iggy.htm https://www.radgametools.com/milesperf.htm
Big thanks for your awesome blog, learnt much from it over the years.
Chance, sure, it's just a matter of logistics. Revision is a bit tricky since it's usually shortly after GDC, a very busy time in the game engine/middleware space I work in, so not usually when I feel up to a pair of international flights. :) Best odds are for something between Christmas and New Year's Eve since that's when I'm usually in Germany visiting family and friends anyway.
I bought my 4790k's ASUS TUF board awhile back because I wanted something basic enough and wasn't interested in overclocking or tweaking. The BIOS had other ideas. I had to manually configure a lot more things just to avoid overclocking, including setting RAM timing and going through each BIOS setting to ensure it wasn't overclocking in some way. The "optimal" setting would turn on aggressive changes like playing with bus speed multipliers, etc.
That or I'm just rewarding shitty corporate product segmentation behavior. I never can quite decide.
I do agree over the recent years getting a "boring" higher-end configuration is getting more and more difficult.
I highly doubt that. XMP is pretty much mandatory to get even remotely close to the intended performance. Without XMP your DDR4 memory isn't going beyond 2400MHz - but you almost have to try to find a motherboard, memory, or CPU which can't run at 3200MHz or even higher. It has all been designed for speeds like that, it's just not part of the official DDR4 spec.
It's less critical with DDR5, but you're still expected to enable it.
https://www.extremetech.com/computing/amds-new-threadripper-...
As much as enthusiasts would like this to be "normalized" - from the perspective of the vendor it is not, they are very clear that this is something they do not cover. And it will become more and more of a problem as generations go forward - electromigration is happening faster and faster (sometimes explosively, in the case of AMD).
But it is quite difficult to get a gamer to understand something when their framerate depends on not understanding it.
https://semiengineering.com/uneven-circuit-aging-becoming-a-...
https://semiengineering.com/3d-ic-reliability-degrades-with-...
https://semiengineering.com/mitigating-electromigration-in-c...
> GD-106: Overclocking AMD processors, including without limitation, altering clock frequencies / multipliers or memory timing / voltage, to operate beyond their stock specifications will void any applicable AMD product warranty, even when such overclocking is enabled via AMD hardware and/or software. This may also void warranties offered by the system manufacturer or retailer. Users assume all risks and liabilities that may arise out of overclocking AMD processors, including, without limitation, failure of or damage to hardware, reduced system performance and/or data loss, corruption or vulnerability.
> GD-112: Overclocking memory will void any applicable AMD product warranty, even if such overclocking is enabled via AMD hardware and/or software. This may also void warranties offered by the system manufacturer or retailer or motherboard vendor. Users assume all risks and liabilities that may arise out of overclocking memory, including, without limitation, failure of or damage to RAM/hardware, reduced system performance and/or data loss, corruption or vulnerability.
It is a scummy little area of technical marketing, it's even in first-party marketing still. I think a lot of ink has been spilled over some real trivial/dumb shit but at least don't lead off showing your product running in an out-of-spec state unless it's clear that's what is doing on.
Fabric overclocking is suuuuper normalized on the AMD side too and it has the same problem. It's higher voltages, and a lot of those "24/7 safe" voltages aren't. On the order of years, under heavy load (not idled down) they do wear out.
I strongly feel like the absolute performance difference is not worth messing around with it anymore. Run ECC at the max supported clock and be done with it. Memory events go into event viewer/dmesg.
Why is that? I’m not a gamer so legit asking. It would seem to me that what would be most important is do the actual games that exist perform well, not some random, hypothetical maximum performance that benchmarks can game.
What's the go-to basic mobo brand/board for non-tweakers these days?
My office glows at night because the RGB dimms stay lit up in sleep mode. But they are fast.
I believe the problems are compounded by the way their SuperIO controls the cooler, because the crashes were associated with temperature excursions to 100C. It's too slow to ramp up and too quick to ramp down. It is possible to tune this from userspace under Linux. But really the up ramp should be controlled by a leading indicator like the voltage regulator instead of a lagging indicator. Alternately the Linux p-state controller could anticipate the power levels and program a higher fan speed.
I have read increasing the ramp up time would smooth out the fan behavior but your experience says this can cause processor fails.
Wow. RAD was bought by Epic? I kinda missed that. Feels old. :(
If that is the case… ouch.
RAD/Epic Games Tools is a small B2B company. Oodle has one person working full-time on it, namely me, and I do coding, build/release engineering, docs, tech support, the works. There's no multiple support tiers or anything like that, all issues go straight into my inbox. Oodle Data in particular is a lossless data compression API and many customers use two entry points total, "compress" and "decompress".
I get a single-digit number of support requests in any given month, most of which is actually covered in the docs and takes me all of 5 minutes to resolve to the customer's satisfaction. The 3-4 actual bug reports I get in any given year, I will investigate.
…it'll need a statement from Intel for some clarity on this…
The motherboard manufacturers are setting default/auto power and current limits that are way outside of Intel's specs (253 W, 307 A) [1].
[0] https://www.tomshardware.com/pc-components/cpus/is-your-inte...
[1] https://www.intel.com/content/www/us/en/content-details/7438... (see pg. 98 and 184, the 13900K/14900K is 8P + 16E 125 W)
> It's not exactly clear why the 13900K suffers from these instability problems, and how exactly downclocking, lowering the power/current limits, and undervolting prevent further crashes. Clearly, something is going wrong with some CPUs. Are they "defective" or merely not capable of running the out of spec settings used by many motherboards?
I'd wager good money on the latter. Why would Intel validate their CPUs against power and current limits that are outside of spec? The users reporting issues probably have CPUs that just made it in to the performance envelope to be binned as a 13900K, so running out of spec settings on these weaker chips results in instability.
It's cases like this where I wish Intel didn't exit the motherboard space, they were known to be reliable but typically at the cost of having a more limited feature set.
Don't guess, measure! The proper action here would be to change BIOS settings from their default / "auto" settings to per-Intel-spec safe ones. Same for RAM, and on systems with known good power supplies, CPU cooling, software installs etc. Then one of the following will happen:
a) BIOS ignores user settings & problem persists.
b) BIOS applies user settings & problem goes away.
c) BIOS applies user settings but problem persists.
Cases a & b count as "faulty BIOS" (motherboard manufacturer caused). Case c counts as "faulty CPU", and replacement cpu may or may not fix that.
No need to guess. Just do the legwork on systems where problem occurs & power supply, RAM, CPU cooling & OS install can be ruled out. Sadly, no doubt there's many systems out there where that last condition doesn't hold.
Still, I don't know if I should RMA. I got the K version because I intended to overclock in the future. And all of this sounds like I won't be able to. I think increasing the voltage a little bit makes the system more stable. I have to play with it. (Really, if someone can say whether I should RMA or not, I would appreciate some input)
Edit: decided to RMA. I have no patience for a CPU that cost me +600€
I don't disagree, but I'm cautious about making a call with the current information available. For example: yes, a "4096W / 4096A" power limit sounds odd, but it's not an automatic conclusion that this limit is intended to work to protect the CPU. Instead, it is a function that allows building a system with a particular PSU dimension — it would be odd if that were overloaded to protect the chip itself. Maybe it is, maybe it isn't.
It's also very much possible that the M/B vendors altered other defaults, but… I don't see information/confirmation on that yet. It used to be that at least one of the settings is the original CPU vendor default, but last I looked at these things was >5 years ago :(.
> It's cases like this where I wish Intel didn't exit the motherboard space,
Full ACK.
Intel 13900K has a fused V/F curve until its maximum Turbo Boost 2.0 (5.5 GHz) in all cores, and two cores at its Thermal Velocity Boost (aka favored cores, 5.8 GHz). How much to boost depends on By Core Turbo Ratio. For stock 13900K, this is 5.8 GHz for 2 cores, and 5.5 GHz for up to 8 cores with E-cores capped at 4.3 GHz.
As you may have noticed, the CPU has a very coarse Turbo Ratio beyond the first 2 cores. This is to allow the clock to be regulated by one of the limits rather than a fixed number. In reality, 253W PL2 can sustain around 5.1 GHz all P-cores, and after 56 seconds it will switch to 125W PL1 which should give it around 4.7 GHz-ish (IIRC).
This is why when a motherboard manufacturer decides to set PL1=PL2=4096 without touching other limits, it results in a higher number in benchmark. The CPU will consume as much power as it can to boost to 5.5 GHz, until it hits one of the other limits (usually 100c TjMax). This is how we ended up in this mess in the consumer market.
Xeon, on the other hand, has a very conservative and granular Turbo Ratio. My Xeon w9-3495x do have a fused All Core Boost that does not exceed PL1 (56 cores 2.9 GHz at 350W), which makes PL2 exist only for AVX512/AVX workload.
(Side note: I always think that PL1=PL2=4096W is dumb since performance gain is marginal at best, and always set PL1=PL2=253W in all machines that I assembled. I think even PL1=PL2=125W makes sense for the most usage. I do overclock my Xeon to sustain PL1=PL2=420W though (this is around 3.6 GHz, which is enough to make it faster than 64-cores Threadripper 5995WX))
Jesus. By German electrical code, you need a 70 mm² cross-section of copper to transfer that kind of current without the cable heating up to a point that it endangers the insulation. How do mainboard manufacturers supply that kind of current without resistive loss from the traces frying everything?
Processors operate at about 1V. At 300W it's enough to use a much smaller cross section, which is split across many traces.
This 350A is flat conductors (maximal surface area thus heat dissipation) and very short (not that much power to dissipate so the things it connects to have a significant effect on heat dissipation).
If you've got 4 layers of 2oz copper, and you make the positive and negative traces 10mm wide, you'll only be dissipating 28 watts when the CPU is dissipating 300 watts. And most motherboards have more than 4 layers and have space for more than 10mm of power trace width. And there's a bunch of forced air cooling, due to that 300 watts of heat the CPU is producing.
Electrical code doesn't let buildings use cables that dissipate 28 watts for 2cm of distance because it would be extremely problematic if your 3m long EV charge cable dissipated 4200 watts.
On the other hand, your national electrical code is going to assume you're running that 350A cable at peak capacity 24/7, right next to other similarly-loaded cables, stuffed in an isolated wall, for very long runs - and it still has to remain at acceptable temperatures during a hot summer day.
The CPU only draws as much power as it needs, though?
I mean, if you plug a 20 watt phone into a 60 watt USB-C power supply, or a 60 watt laptop into a 100 watt USB-C power supply the device doesn't get overloaded with power. It draws no more current than it needs.
The motherboard's power limits should state the amount of power the PCB traces and buck regulators are rated to provide to the socket - and if that's more than the processor needs that's good, as it avoids throttling.
Of course users, especially enthusiast motherboard consumers, hate throttling, hence the default.
* CPU clock speed increases (turbo boost), as long as the CPU isn't hitting: 1) Tj_MAX (max temp before thermal throttling kicks in); 2) the power and current limits specified by the motherboard (in this case, effectively disabled by the out of spec settings).
* Weaker chips will require more power to hit or maintain a given turbo clock speed: with the power and current limits disabled, the CPU will attempt to draw out of spec power & current, causing issues for the on die fully-integrated voltage regulator (noting that there's also performance/quality variance for the FIVR), resulting in the user experiencing instability.
Some consider it cheating the benchmarks, but the justification is that TDP is the Thermal Design Power. It's about the cooling system you need, not the power delivery. If you make reasonable assumptions about the thermal inertia of the cooling system you can Turbo Boost at higher power and hope the workload is over before you are forced to throttle down again.
Any mainboard that sets power limits to the TDP would be considered wrong by both the community and Intel. This looks like a solid indication that the issue is with Intel
Boss says, "Do the thing." Engineer says, "The thing is out of spec!" Boss says, "Competitor is doing the thing already and it works." Engineer does the thing.
people forget, blowing up AM5 cpus wasn't just as Asus thing... they were just the most ham-handed with the voltages. Everyone was operating out of spec, there were chips that blew up on MSI and Gigabyte boards, and it wasn't just X3D either.
Intel is no different - nobody enforces the power limit out of the box, and XMP will happily punch voltages up to levels that result in eventual degradation/electromigration of processors (on the order of years). Every enthusiast knows that CPU failures are "rare" and yet either has had some, or knows someone who's had some in their immediate circles. Because XMP actually has caused non-trivial degradation even on most DDR4 platforms.
In fact it's entirely possible that this is an electromigration issue right here too - notice how this affects some 13700Ks and 13900Ks too? Those chips have been run for a year or two now. And if the processors were marginal to begin with, and operated at out-of-spec voltages (intentionally or not)... they could be starting to wear out a little bit under the heaviest loads. Or the memory controllers could be starting to lose stability at the highest clocks (under the heaviest loads). That's a thing that's not uncommon on 7nm and 5nm tier nodes.
Epic Games Tools is B2B and we don't generally get bug reports from end users (although later last year, we did have 2 end users write to us directly because of this problem - first time this has happened for Oodle that I can think of, and I've been working on this project since 2015). Point being, we're normally at least one level removed from end user bug reports, so add at least a few weeks while our customers get bug reports from end users but haven't seen enough of them yet to get in touch with us (this is a rare failure that only affects a small fraction of machines).
13900Ks have been out since late Oct 2022. It's possible that this doesn't show up on parts right out of the box and takes a few months. It's equally plausible that it's been happening for some people for as long as they've had those CPUs, and the first such customers just bought their new machines late 2022, maybe reported a bug around the holidays/EOY that nobody looked at until January, and then it took another 2-3 months for 3-4 other similar crashes to show up that ultimately resulted in this case getting escalated to us.
For the majority of systems "in the wild", I don't know. We had two people with affected machines contact us and consent to do some testing for us, and in both cases the issue still reproduced after resetting the BIOS settings to defaults.
No shielding, earth - using the crappiest/cheapest PC they could get instead of using the recommended kit as the sales droid wanted a bigger commission.
Said call me when you replaced the h/w - I walked out and went to the airport. They never called me.
DDR4/Intel motherboards are cheaper than AM5/DDR5 - also a Ryzen laptop foobarred on my daughter so to me Intel kit was just more stable - no weird XMP issues or overclocking to the nines.
For the situation of an actual (consistent) hardware bug, redundancy wouldn't help… the redundant system would have the same bug. Redundancy only helps for random-style issues. (Which, to be fair, the one we're talking about here seems to be.)
It can also misbehave without any hardware bugs due to glitching. Rates of incidence of this must be quite low or that would be considered a HW bug, but it's never zero. Run code for enough hours on enough machines collecting stack traces or core dumps on crashes and you will notice that there's a low base rate of failures that make absolutely no sense. (E.g. a null pointer dereference literally right after a successful non-null pointer check 2 instructions above it in the disassembly.)
You will also notice that many machines in a big fleet that log such errors do so exactly once and never again, but some reoccur several times and have a noticeably elevated failure rate even though they're running the exact same code as everyone else. This too is normal. These machines are, due to manufacturing variation on the CPU, RAM, or whatever, much glitchier than the baseline. Once you've identified such a machine, you will want to replace it before it causes any persistent data corruption, not just transient crashes or glitches.
Sure, the user has a broken CPU, but if you can work around it and still let the user play their games, you should.
Software running at scale in the cloud is written to be resilient to errors of this nature; jobs are cattle, if jobs get stuck or fail they are retried, and sometimes duplicate jobs are started concurrently to finish the entire batch earlier.
Consumer machines are comparably wild. Remember that this issue was mainly spotted from Unreal error messages. Some do too much overclocking without enough testing, which will eventually harm the hardware anyway. Some happen to live in places where single-error upsets are more frequent (for example, high altitude or more radioactive bedrock). Some have an insufficient power supply that causes erratic behaviors only on heavy load. All those conditions can be handled in principle, but are much harder to do so in practice. So giving up is much more reasonable in this context.
It sounds like a hardware issue, i'm guessing over-agressive memory/cpu tuning, underpowered PSU triggering off behaviour etc. The fact that replacing the processor makes the problem go away does not in itself point to the processor as the issue - you may find that changing the memory also 'fixes' the problem.
Sure it's unpleasant to be the messenger of bad news, but the alternatives are far worse unless the system is just a dedicated game console without any background processes (which isn't how those CPUs are used).
Was a real pain to deal with the fallout...
The ball is in Intel's court, such faulty CPUs should never have made it out into the wild.
However, the really useful thing that this software could do is make the error message much better - explaining the likely thing that is causing the decompression to fail, and advising the computer should be fixed.
Oodle itself has some diagnostics hooked up to logs but none of that is user-facing. All the user-facing stuff needs to get handed off multiple times to get from low-level IO plumbing to somewhere that even knows how to display a user-facing error message to begin with.
The most common cause for that error message was, and continues to be (except on the relatively small number of affected machines), that compressed shader data on disk is corrupted. If and when we have a handle as to what actually causes the problem and a minimally-invasive fix (as opposed to the list of several different anecdotal "what if you try X?" that we got from Intel HW lab folks), we'll try to detect affected machines and point them to a website with instructions. For now, it's just a random error message that happens to frequently show up on machines encountering this issue, and all that would change if we changed the error message was that we'd confuse end users more and have to list two different error messages on that page instead of one.
Then the natural question arises: if we detect that it is a single bit flip, should we "un-flip" that bit, fix the data, and continue? The answer is: no. These types of errors should be explicit. They successfully help to detect broken RAM and broken network devices, that have to be replaced. However, the error is fixed automatically anyway by downloading the data from a replica.
A relatively major realization during the investigation was that a different mystery bug that also seemed to be affecting many Unreal Engine games, namely a spurious "out of video memory" error reported by the graphics driver, seemed to be occurring not just on similar hardware, but in fact the exact same machines.
For a public example, if you google for "gamerevolution the finals crash on launch" and "gamerevolution the finals out of video memory", you'll find a pair of articles describing different errors, one resulting from an Oodle decompression error, and one from the graphics driver spuriously reporting out-of-memory errors, both posted on the same day with the same suggested fix (lower P-core max clock multiplier).
That's the problem right there in a nutshell. It's not just Oodle detecting spurious errors during its validation. Other code on the same machine is glitching too. And "just try repeating" is not a great fix because we can't trust the "should we repeat?" check any more on that machine than we can trust any of the other consistency checks that we already know are spuriously failing at a high rate.
Many known HW issues you can work around in software just fine, but frequent spurious CPU errors don't fall into that category.
There are two after market contact frames that drop the temperature around 10 Celsius and ensure flat contact with the head spreader. The stock frame causes the center of the head spreader to dip.
I wonder if the turbo boost is controlled by a Proportional–integral–derivative controller.(PID)
The idea that the parameters are fine tuned to slow down processor speed as it heats up but before it overshoots its maximum threshold.
If those PID values are tuned to assume flat heat spreader/heat sink contact, I can see where a bent heat spreader could cause the cpu to overshoot its safe limit and cause errors.
> This has resulted in hundreds of CPUs detected for these errors
I don't find this situation much different than needing to dial back BIOS settings when actual crashes are observed.
Not a good look but at least it’s fixable with bios tweaks rather than a silicon flaw that’s permanent
First time I see this nice euphemism for overclocking.
A pattern I've noticed is that some of the AMD systems I administer have never crashed or panicked. Several are almost ten years old and have had years of continuous uptime. Some have had panics that've been related to failing hardware (bad memory, storage, power supply), but none has become unstable without the underlying cause eventually being discovered.
Intel systems, on the other hand, have had panics that just have had no explanation, have had no underlying hardware failures, and have had no discernible patterns. Multiple systems, running an OS and software that was bit-for-bit identical to what has been running on AMD systems, have panicked. Whereas some of the AMD systems that had bad memory had consumer motherboards with non-ECC memory, the Intel systems have typically been Supermicro or Dell "server" systems with ECC.
In one case two identical Supermicro Xeon D systems with ECC were paired with two identical Steamroller (pre-Ryzen) AMD systems. All systems provided primary and backup NAT, routing, firewalling, DNS, et cetera. The Xeon systems were put in place after the AMD systems because certain people wanted "server grade" hardware, which is understandable, and low power AMD server systems weren't a thing in that time period. Over the course of several years, the Xeon systems had random panics, whereas one of the AMD systems had a failed SSD, but no unplanned or unexplained panic or outage, and the other had never had a panic or unplanned reboot in all the years it was in continuous service.
Had I collected information more deliberately from the very beginning of these side-by-side AMD and Intel installations, I'd have something more than anecdotal, but I'm comfortable calling the conclusion real: multiple generations of Intel systems, even with server hardware and ECC, have issues with random crashes and panics, on the order of perhaps one every year or two. I do not see a similar instability on AMD, though.
With brand new Intel CPUs taking substantially more power than similarly performing AMD CPUs, we have a more literal example of what I think is the underlying cause: Intel is trying way too hard to get every tiny bit of performance out of their CPUs, often to the detriment of the overall balance of the system. Between the not insignificantly higher number of CPU vulnerabilities on Intel due to shortcuts illustrated by the performance losses from enabling mitigations, and the rather shocking power draw of stock Intel CPUs that have turbo boosting enabled, I can't recommend any Intel system for any use where stability matters.
It could keep running for 70 hours, or it could crash twice in 4 hours. Stress-testing CPU, GPU, memory, and storage doesn't invoke a crash, but it'll crash when all I'm running is a single Firefox tab with HN open.
Maybe I got unlucky, or maybe you got lucky. Who knows, really.
On most computers, holding the power button for five seconds turns off the power, whereas a hard reboot might be similar to someone pressing the reset button. If it were me, I'd try, in order, checking that you have the correct version of your BIOS (I accidentally loaded the latest BIOS on a system with a Ryzen 2600, and it made memory transfers unhappy), using another power supply for a while, running another OS, like NetBSD, on an external drive and doing lots of test compiles, clocking the RAM one or two steps slower, then removing all but one DIMM at a time and testing with just one DIMM.
I had this actually with my Silverstone cased HTPC. Reboots several times shortly after bootup. Random. Turns out the power button was actually defective. Who would have guessed that.
On my Asus/14900k, it was uncapping PL1/2 and I saw absurd temps and power every time anything even touched the CPU. I programmed PL1/2 to 125/253w per Intel ARK and everything normalized.
I did not do Prime95 at the insane default power limits but I suspect similar.
> For MSI:
> Solution A): In BIOS, select "OC", select "CPU Core Voltage Mode", select "Offset Mode", select "+(By PWM)", adjust the voltage until the system is stable, recommend not to exceed 0.025V for a single increase.
This really sounds like the Intel defaults are broken too.
Per the article MSI literally suggests to OC the CPU to fix the problem