I don't think any programmer puts UB on purpose into the code, even if they can enumerate from memory all 200 or so cases of UB just in the C standard (no idea how big the list is in the C++ standard - thousands maybe?).
The original sin was compiler writers exploiting UB for optimisations in bizarre ways instead of working with the C and C++ committees to fix the language to enable those types of optimizations without requiring UB, or at least to classify UB into different 'hazard categories' (most types of UB in the standard are completely irrelevant for code generation and optimization)
I agree with you. But the problem is not people putting UB in their code, either on purpose or by mistake: we all do that, every day!
The problem is people trying to defend their code once it has been made clear to them that it contains UB, and trying to fight the compiler rather than fix their error.
How about when code is written correctly and then later the standards body makes previously implementation defined behaviour into undefined? I've got some code that calls realloc that WG14 declared to be undefined years after I wrote it and I doubt that experience is unique.
The evangelical attitude that the standard committee knows best and your code is wrong and you should immediately down tools to work around the compiler noticing the opportunity to miscompile it is quite popular. I think it's a really compelling argument to build nothing whatsoever on the ISO C/C++ stack and replace what you do have with something less hostile, aka anything whatsoever - rust, python, raw machine code written in hex - none of them have this active hostility towards the dev baked into the language design.
For realloc, different implementations did different things and clearly said they will not change. There wasn't really any other choice. If your program was written for one implementation where it works, it can continue to do so, but it was never portable to other implementations. The standard now simply reflects this reality.
Thank you for pointing this out.
Put the two together and you get a fast and fragile language implementation. I know why the benchmark people push the compiler in that direction. I'm doubtful that WG21 or WG14 especially want this emergent property.
My suspicion is that this is an accident of history that has too much unwarranted inertia behind it. The moral stance that it's all lesser programmers erroneously writing wrong code is aggravating in that context as it actively opposes anyone making things better.
Compare to the concurrent memory model: while DRF-SC still has a plenty of UB, it is at least possible for a competent programmer to figure out the correctness of their code.
I certainly do not believe that is realistic to stamp out all UB from C and C++ (at least while pretending that the resulting languages have anything to do with the original ones), but there is a lot that the standard could do to try to limit the most egregious cases, possibly providing different levels of conformance (like it is done for floats and IEE754).
[1] of course implementors are part of the committee so they are not blameless.
I don't know how constraining is this more restrictive implementation of the standard is, but certainly it will help with maintaining a bit of sanity.
No, they cannot. You don't have the right to make any assumption about integer overflow in C.
Most of the whining about UB is from people who still refuse to accept that you don't have the right to think about your processor family once you write in any language that is not assembly.
Since when is it reasonable to assume that?
> Then you have weird edge cases like assigning the return value of a two argument std::max involving temporaries to a reference
You have a reference to a temporary. Reference lifetime extension is a thing. No UB there. Completely defined and supported.
However, some benchmarks use iteration on a signed integer, and assuming that loop terminates makes it slightly faster, so in order to retain that marginal advantage over other languages, signed iteration shall be assumed to never overflow.
This is very typical of the C++ experience.
I'd actually go the other way: I think most practicing programmers have not read the entire standard specifying most languages they use day-to-day and have no real idea what the abstraction-break looks like that turns their code into a format consumable by the next layer down.
How many Java programmers do we assume know anything about the bytecode of the JVM, for example?
The corresponding Java, C#, Python standards, equally with the standard library, are even bigger.
To come back to the point, many don't know how deceptively complex Python happens to be, even though on the surface looks like a BASIC replacement.
If we compare to equivalent Python code, for example, the behavior is simple and straightforward: division by zero causes a runtime error that is reported in a well-defined manner. Similarly with other arithmetic issues in C++ that can trigger UB - e.g. integer overflow is just not a thing in Python (short of OOM). The detailed rules may well be complicated, but it doesn't matter as much when the behavior is intuitive and conforms to common sense expectations.
Further, those 'checks' might be in place for a long time, outlasting compiler versions and perhaps even language standards.
Indeed, it isn't. The check becomes useless if it happens after the division, though.
As a slightly contrived example, assert(sizeof(char) == 1) is true by definition and could be elided, but it might be a useful reminder to see it in the source, next to code that implicitly relies on this truth.