> And UB shouldn't be a scary thing 99% of the time, as you won't be hitting it anyway (or its actually the result of another real bug, like not handling overflows etc.), though as at some level you start trying on platform specific behavior and start /defining/ them
I’m talking about language-level UB - that’s not anything you can rely on on any toolchain / platform. And UB will also manifest as random unexplainable crashes in random spots just like memory corruption will.
As for tooling, unless you’re using something 100% memory safe with absolutely no call out to unsafe code (which definitely isn’t a game/GPU scenario), you’re going to have this risk & it almost doesn’t matter how much you test or use memory checkers because the long tail of issues are going to be the result of a really difficult to reproduce sequence of events. Additionally, if you have any multi-threaded code, all your testing goes out the window because I’ve seen many concurrency bugs that have hidden in plain sight in very public spots until someone figured out a repro (e.g. https://reviews.llvm.org/D114119, https://probablydance.com/2022/09/17/finding-the-second-bug-...). Race conditions are notoriously difficult to reproduce in controlled environments.
Oh and when I said you’re safe if you’re in a 100% memory safe language? I lied. There are compiler bugs that could be misgenerating your code + you’re likely running unsafe constructs somewhere in your code that has access to your memory space (whether the OS kernel, whatever is making syscalls to the kernel on your behalf, something that eventually calls the C library, something that is needed for performance, etc etc etc) not to mention bugs in the compiler/JIT that either allow unsafe/unsound constructs or just straight up miscompile correct code. And finally, there are sources of HW issues unrelated to memory bitflips (your CPU is super complex and has bugs too you know as does your memory controller).
> research has shown that the majority of one-off soft errors in DRAM chips occur as a result of background radiation, chiefly neutrons from cosmic ray secondaries,
> Recent studies[5] show that single-event upsets due to cosmic radiation have been dropping dramatically with process geometry and previous concerns over increasing bit cell error rates are unfounded.
So aside from the fact that DRAM susceptibility to cosmic rays has been decreasing (I’m not convinced with the Wikipedia explanation - an alternate explanation can be that the percentage of critically important data in DRAM has shrunk as a percentage as the overall capacity has increased), you’re argument would be that random cosmic radiation is going to randomly hit the DRAM cell containing your code / critical data. On the order of things that are likely, that’s the last thing.
Oh and all this relies on you correctly grouping related crashes correctly and I’ve generally seen that to be a significant challenge on any project I’ve participated on (e.g. Oculus for the longest time would group unrelated crash reports and not group the same ones correctly although I helped the team try to make progress on that).
Again, it’s not impossible that it’s a legit HW corruption issue. However, all the engineers I know frequently also blamed cosmic rays but at the end it’s all just shorthand for “not worth wasting time trying to track down because it’s in the tail of issues you’ll never get to“.