This is because of the particulars of the C language. For instance, the way it handles arrays and pointer arithmetic make it difficult to detect every instance of out-of-bounds access.
[0] -fwrapv, see https://gcc.gnu.org/onlinedocs/gcc/Code-Gen-Options.html
Using out of bounds references in arrays and pointer arithmetic are unsafe but they can be defined. The definition of "array accesses using indexes that are outside the range of 0 - the declared array size will access memory in the data segment. The compiler will treat that access as if the addressed memory had the type specified for the array." The behavior is defined, it is unsafe, but it is defined. Not like "the compiler may or may not choose to optimize out loads and stores, or re-order them." which especially on embedded systems creates bugs where the things like "read the status register THEN read the data register (which clears the status register when read)" those get re-ordered and suddenly your loop never exits because your status bit is never set. That kind of UB needs to die in a fire IMHO.
Why are the embedded systems programmers not using `volatile`, which exists for this exact reason?
All it does is forbid reordering or removing accesses to a particular memory location.
Historically, many compilers implemented it as a hard memory barrier, but that isn't how the standard defines it.
The compiler is free to reorder accesses to multiple volatile variables if they happen prior to the same sequence point. So (roughly) expressions involving two volatile variables do not have those accesses sequenced.
But you are right that the usage described earlier is OK, as long as both variables are marked volatile, and the accesses straddle a sequence point.
The more common mistake with volatiles is to use the for multithreading primitives.
C lacks the intrinsics you'd actually want for this (explicit load/ store of various machine sizes) but it does provide the "volatile" storage qualifier which is what you should be using to do what you apparently wanted in C today. It makes no sense to demand everybody else writes some sort of "not-volatile" qualifier in front of every variable to tell the compiler that actually this is just a variable and it's OK to optimise.
I believe performance (compiler optimisation) is the reason the language isn't defined that way. Permitting the compiler to assume that the runtime error will never arise, opens the door to all sorts of optimisations. (At least, that's the idea.)
C permits the 'union trick' to (roughly speaking) access the bit-pattern of a value as another type, which is to say an escape-hatch is offered in the language.
Similarly the strict aliasing rule is surprising to people who are new to C, but the C standard committee seem to be committed to keeping it, presumably for performance reasons.
> Not like "the compiler may or may not choose to optimize out loads and stores, or re-order them." which especially on embedded systems creates bugs
Right, but it's defined that way to enable compiler optimizations, not to spite the programmer. As others have mentioned, C has features like volatile specifically to address this kind of thing. If the C standard required memory fences to be inserted everywhere, performance would be ruined.
> That kind of UB needs to die in a fire IMHO.
C cannot easily be made into a safe language, and I think the committee is doing the right thing in declining to try to make C into something it isn't. On the plus side, there are plenty of other languages around, many of them with compelling advantages over C. Ada, Rust, and Zig, for instance.
Then again, just use C++ alongside std::vector, std::array and std::string with FORTIFY turned on (or equivalent) and be done with it.
If you do pointer arithmetic to derive a pointer value pointing 2 or more elements beyond the final element of an array, that's undefined behaviour, even if you never dereference that pointer.
But the current trade-off (performance always wins) means that kernels and other embedded-style programs cannot rely on the compiler doing the "reasonable" thing for UB because it's explicitly allowed to do whatever it wants (which is generally, try for better performance).
For kernel-style work, slightly lower performance but predictable/defined behavior for some of what is currently UB, makes life much simpler.