> The compiler doesn't need to error on all instances of this, but a flag to have the compiler warn or error if it would remove a unique branch and that branch may exit prior to returning control, that could trigger the desired behavior. If the compiler has enough knowledge to determine redundant code, then it has enough knowledge to know whether it is removing code that is not due to duplication.
I wrote "redundant", not "duplicate", intentionally. A check is redundant if other code implies that it cannot possibly ever end up true. It being a duplicate check is not the only way for that to happen, and it's actually the exception. Also, no, the compiler most likely doesn't have that knowledge. A compiler doesn't work the way you seem to think.
> Any new language introduced today that said "well, there's some interesting interactions sometimes if you don't pay close attention, and the compiler/VM might remove statements you write because they are testing those same weird interactions[1] as it things they can't happen" would be laughed out of town.
Or, more likely, it wouldn't. These are good reasons to not write high-level software in C, but there are also good reasons why people who actually need high speed still do use C. And it's not that we like the fact that writing correct C is hard.
You seem to imply that there is no real reason for C behaving the way it does, and that those rules for undefined behaviour only exist to make the life of programmers miserable. Those things are undefined because making them defined would actually be expensive in terms of performance.
> That the optimizations that C compilers do are complex and have many stages of optimization is not a suitable counter for the criticism that those same optimizations sometimes cause non-obvious interactions with safety checks meant to test the same edge cases that those optimizations take advantage of.
The way you phrase things suggests that you might be confused about how the compiler "reasons". The compiler doesn't read your program, sees what you mean, and then tries to find loopholes in order to misunderstand you. The compiler reads your program, and only understands what your program means according to the formal specification of the language that you claim it is written in. In the example I gave above, you seem to think that there is a call to abort() that the compiler "removes". There isn't. According to the C spec, that call is unreachable, and as such the semantics of that statement is a noop, which is what the compiler will correctly map to machine code, somehow. Explicit dead code removal on some intermediate representation of the AST is just one way that could happen, and it's an implementation detail of the compiler.
> That's like someone saying for safety reasons you need to see at least 20 feet of road in front of you at all times per 10 MPH on the highway, and people complaining about how that's not feasible because it would force you to slow down 10-20 MPH occasionally as you went around some turns. Yes. Yes it would.
Yep, that's a perfect analogy for people complaining that when they talk to a C compiler, they maybe should be writing C, and not a language that they themselves made up, if they expect the compiler to understand them. Even it's sometimes difficult.
> Just because you can do an optimization, doesn't mean you should.
I agree. But for the most part, that's not what's happening. For the most part, optimizations are not intended to break your code, but rather it so happens that, in order to optimize some code, you have to rely on all code having certain correctness properties (that it should have if it is C code, according to the C spec), which then, unfortunately, happens to break some code that doesn't have those properties. Often it's somewhere between infeasible and impossible to distinguish those cases that are correct and can thus be optimized without introducing unwanted behaviour from those that are not correct and thus break as a result.
> I'm not the original commenter, but if you're referring to my suggestion of at least warning when entire unique branches of the original code are removed, that's entirely possible. That it might require reworking or ever removing portions of the current optimization pipelines, or cause compilation speed to slow considerably is irrelevant to this particular aspect, because right now we aren't having a discussion of whether it's worth it, but whether it's event possible.
First, equivalence between pieces of code is undecidable, and second, "unique branches" is not in any way a useful concept anyway.
Sure, you can try to make a compiler detect certain instances of what seems like a safety check that doesn't ever trigger. But that will either be completely ineffective (as it only detects a small minority of cases), or it will produce tons of bogus warnings (because there are tons of cases where it is perfectly sensible to have "unique branches" in your code that are provably never taken, even ones that abort the program, and it is logically impossible to distinguish those from ones that were written with the intent to catch a runtime exception that the programmer expects to actually happen at runtime).