Sure, the user has a broken CPU, but if you can work around it and still let the user play their games, you should.
Sure, the user has a broken CPU, but if you can work around it and still let the user play their games, you should.
Sure it's unpleasant to be the messenger of bad news, but the alternatives are far worse unless the system is just a dedicated game console without any background processes (which isn't how those CPUs are used).
Was a real pain to deal with the fallout...
Software running at scale in the cloud is written to be resilient to errors of this nature; jobs are cattle, if jobs get stuck or fail they are retried, and sometimes duplicate jobs are started concurrently to finish the entire batch earlier.
Consumer machines are comparably wild. Remember that this issue was mainly spotted from Unreal error messages. Some do too much overclocking without enough testing, which will eventually harm the hardware anyway. Some happen to live in places where single-error upsets are more frequent (for example, high altitude or more radioactive bedrock). Some have an insufficient power supply that causes erratic behaviors only on heavy load. All those conditions can be handled in principle, but are much harder to do so in practice. So giving up is much more reasonable in this context.
A relatively major realization during the investigation was that a different mystery bug that also seemed to be affecting many Unreal Engine games, namely a spurious "out of video memory" error reported by the graphics driver, seemed to be occurring not just on similar hardware, but in fact the exact same machines.
For a public example, if you google for "gamerevolution the finals crash on launch" and "gamerevolution the finals out of video memory", you'll find a pair of articles describing different errors, one resulting from an Oodle decompression error, and one from the graphics driver spuriously reporting out-of-memory errors, both posted on the same day with the same suggested fix (lower P-core max clock multiplier).
That's the problem right there in a nutshell. It's not just Oodle detecting spurious errors during its validation. Other code on the same machine is glitching too. And "just try repeating" is not a great fix because we can't trust the "should we repeat?" check any more on that machine than we can trust any of the other consistency checks that we already know are spuriously failing at a high rate.
Many known HW issues you can work around in software just fine, but frequent spurious CPU errors don't fall into that category.
The ball is in Intel's court, such faulty CPUs should never have made it out into the wild.
Then the natural question arises: if we detect that it is a single bit flip, should we "un-flip" that bit, fix the data, and continue? The answer is: no. These types of errors should be explicit. They successfully help to detect broken RAM and broken network devices, that have to be replaced. However, the error is fixed automatically anyway by downloading the data from a replica.
It sounds like a hardware issue, i'm guessing over-agressive memory/cpu tuning, underpowered PSU triggering off behaviour etc. The fact that replacing the processor makes the problem go away does not in itself point to the processor as the issue - you may find that changing the memory also 'fixes' the problem.
However, the really useful thing that this software could do is make the error message much better - explaining the likely thing that is causing the decompression to fail, and advising the computer should be fixed.
Oodle itself has some diagnostics hooked up to logs but none of that is user-facing. All the user-facing stuff needs to get handed off multiple times to get from low-level IO plumbing to somewhere that even knows how to display a user-facing error message to begin with.
The most common cause for that error message was, and continues to be (except on the relatively small number of affected machines), that compressed shader data on disk is corrupted. If and when we have a handle as to what actually causes the problem and a minimally-invasive fix (as opposed to the list of several different anecdotal "what if you try X?" that we got from Intel HW lab folks), we'll try to detect affected machines and point them to a website with instructions. For now, it's just a random error message that happens to frequently show up on machines encountering this issue, and all that would change if we changed the error message was that we'd confuse end users more and have to list two different error messages on that page instead of one.