The Value of Undefined Behavior
nullprogram.com
nullprogram.com
If they want wrapping overflow, fine. If they don't, great. Tell the compiler don't make it guess/assume/optimize/hope. What this article is arguing for IMO is not the concept of behavior that happens to sometimes work and sometimes not depending on the target machine (because thats crazy) but the functionality itself, which is valuable, and should be explicit and defined. The issue is the last-generation languages like C/C++ don't give you mechanisms to express what you want so you're left hoping the compiler figures out what you meant. That's not a world I want to live in anymore, we can do better.
UB is bad, period. Define what you want, make the compiler do it.
So the attitude of "only unsafe programs have UB by the Rust spec" effectively means "all programs have UB by the Rust spec"
I don't think that's a particularly useful point, as is.
I have no idea whether any equivalent exists for Rust.
One can also forbid the use of any `unsafe` blocks for a crate/module (with the #[forbid(unsafe_code)] attribute), but this is slightly different to the {-# LANGUAGE Safe #-} pragma in that it does consider imports at all. However, since "unsafety" is modeled in the language in Rust, and is granular, people are encouraged to have every module be "Trustworthy" (in Safe Haskell terms): if an individual function could cause undefined behaviour if misused, it should be marked "unsafe", and a "Safe" module (i.e. one using the forbid above) won't be able to use it.
One reason for that is that it's easy to get rid of it, and guess what, we have a bunch of compiler flags just for that.
That absolutely does not mean that "UB is bad, period".
In fact you praise Rust but only safe Rust has no UB.
Actually, if they do end up putting more and more knowledge about the type system invariants in the optimizer, the rules for UB in unsafe code are going to be more complicated than C (think type-based aliasing on steroids).
Don't get me wrong, it's good design to nicely package these in unsafe code, but if your language can craft pointers and load/store from them, your spec will have to have UB (or boundcheck all loads and stores).
This is one of the biggest issues I have with Rust, personally. The interaction between safe and unsafe Rust is not well defined, making it largely impossible to write 'correct' unsafe Rust that isn't broken in some weird situation, or won't ever get broken in later versions of the compiler. And since the Rust compiler has even more information then the C compiler does, it can theoretically make much crazier optimizations then it is doing right now that could break your code down the line. Like you mentioned, if it goes down that route it'll be like the strict-aliasing situation in C but times 10.
Or from a more pragmatic point of view, if one could compile un-safe Rust similarly to "optimisations off" in C compilers. No aggressive optimisation, only straight forward translation. While Rust code marked safe could be compiled with the most insane optimisations, and still be ... safe.
If not, I think people concerned about such things, might opt to write the unsafe parts on C (a known quantity) and the safe parts in Rust.
All this stuff is why we’re working so hard on formalizing unsafe! Your parent is right that it’s a bit fuzzy right now, but in the future, we hope that it’ll be way nicer than other unsafe languages. There’s some tooling ideas...
For the curious, the performance uses of unsafe in encoding_rs are roughly in these categories:
- Omitting Unicode correctness checks (i.e. asserting that output is correct without having the standard library re-check; Unicode correctness is part of the core competence of the crate).
- Viewing buffers of u8 or u16 as buffers of usize.
- Viewing buffers of u8 or u16 as SIMD register-sized units.
- Omitting bound checks where the compiler fails to elide them and it matters for performance, especially where the performance difference matters for competitiveness with C++.
- Reinterpreting SIMD registers in a different lane configuration. (Exposes endianness but otherwise actually safe.)
- Calling intrinsics that don't really have anything unsafe about them except that Rust makes intrinsics categorically unsafe instead of deciding on a case by case basis if they need to be unsafe.
Of these, viewing buffers as buffers of ALU or SIMD register-sized units are the only cases that come near strict aliasing and could be trouble if the compiler felt eligible to reorder writes of (in the C sense) incompatible types relative to each other on the assumption that buffers of different types are disjoint.
Not contesting the claim, but what in the interaction specifically is not well defined?
I don't see why it's so bad. I've never really seen a legitimate use for signed overflow, and if it happens, the code is likely buggy anyway.
"I don't know how to properly check overflows so I'm just going to look and see if the number has wrapped" isn't really a good excuse, even if a few people fall for that trap. If anything, I'd rather see a simple, standard way for checking arithmetic introduced.
That's way beyond "can't do X"; it's "the meaning of your program suddenly ceases to exist".
If they want wrapping overflow, fine. If they don't, great.
What if I know overflow can't happen (e.g. I have an integer counting from 1 to 10)? How do I say "it doesn't matter how you handle overflow; just do whatever is fastest"?
In the future we’ll have const generics that let you express it directly without that shenanigans.
type jump_by_tenth is delta 0.1 range 1.0 .. 99.0;ETA: I'm not completely correct. Ada could have used something akin to 2-complement overflow handling, that is forcing every operation on a range to have a result in the same range, but it introduces runtime overhead too.
* overflow is a “program error”, not UB
* compilers must either check for overflow and panic, or two’s compliment wrap
* when debug_assertions are on, implementations are required to panic on overflow
This means release mode wraps today, but if it’s cheap enough someday, we could make it panic.
If you want specific semantics on overflow, then you should use the various wrappers and/or methods that let you do them directly, rather than relying on any of the above. But you don’t have to specify; by default, the above is the semantics.
This argues that it is precisely C's undefined behavior which allowed for its dominance.
In fact, the ANSI C spec even states from your link that "Undefined behavior gives the implementer license not to catch certain program errors that are difficult to diagnose." That's just not an acceptable tradeoff anymore -- in fact I don't know it ever was. That's akin to shipping your org chart because you can't make your program do what everyone agrees it should do. You know, catch programming errors, in this case.
Your program is still doing what you told it to if you weren't relying on any undefined behavior (Which, obviously, you shouldn't be). That's the point - if you write C that doesn't rely on any UB, then it can run on a large variety of architectures, work exactly the same on each one, and still get close to the maximum speed for each architecture because they can take the fastest possible option for any of the UB cases (And your program shouldn't care what they do in that situation).
When you add type declarations and (optimize (safety 0)) into the mix you get similar scope of UB as in C (another thing is that in reality nobody actually writes code like that)
Programs are not all written by experts. Performance may be left on the table in areas that have nothing to do with optimizers or schedulers (e.g. picking an absolutely terrible algorithm or not noticing unnecessary memory allocations).
If something must remain “unspecified”, I at least want a debugging tool for that behavior. If something “may” happen, give me a switch to make it happen. If something is data-dependent, show me the data that triggers a different outcome or give me a way to scramble data enough to increase the chance of triggering a change in behavior.
Then you come up with the "solution" that everybody must always check for NULL pointers before dereferencing them. If everybody always did that we might as well just define the behavior of reference-after-free, because what you said either a) requires you to keep a reference count (or null-check before every deref which may be even more expensive) at runtime(!), or b) requires a program-level proof.
So... no. What you suggested is not a solution in any meaningful sense.
If you never reuse addresses, long running programs will have some trouble ahead.
Oh, it's not even that simple because you'd have to make access to the region atomic.
There's certain classes of C's UB that can be addressed, but most of it can never be without neutering C to the point of uselessness.
I think better diagnostics for this would be a win for everyone.
Ya, I know, offtopic.
* With integer overflow it would be relatively easy to write a spec that allows "n+1>n" optimizations but not utterly anything to happen. With strict aliasing it's probably not too hard either. You'd still see some program misbehavior in the strict aliasing case, but it would be a vastly reduced set of misbehavior. And allowing signed numbers to act more like abstract numbers would often decrease misbehavior.
Magic behavior of "int" (but not "unsigned int") --> just use a type that's the correct size and better conveys your intention, like "size_t".
(And that one doesn't make your code "many times faster", in the example given it saves a single instruction.)
Type-based assumptions about aliasing (with a special case for char*) --> be explicit about your intention with the "restrict" keyword.
Saving one instruction was probably worth it in the 70s, when security wasn't a thing, but it definitely isn't now.
And please, if the compiler can prove that the program invokes UB in "strict-UB" mode, that should be an error, not a cause for the code to be deleted. I'm talking about the case of this SPEC 2006 code [1]
It doesn't make sense to me to just apply potentially disastrous micro-optimizations that save a couple of percent of run-time on the whole program. This sounds like a classic case of optimizing before measuring.
But you have done the measurements, right?
Also compilers normally do not prove that programs hit UB, quite the contrary. Doing that statically is, in general, impossible.