Does Rust define what I get when I dereference NULL in unsafe code? I doubt it, since it would require a NULL check before every pointer dereference.
The only insane thing about UB is that compiler writers took what everyone understand meant "the compiler emits what it emits and you get what you get" and turned it into "since it's undefined it means it can never happen so we can delete your null check".
(Here's a link with a bit more detail: https://courses.cs.vt.edu/cs3214/spring2026/questions/catchi...)
So, while it's accurate to say that inserting null checks before every dereference is one way that you could implement this to make it well-defined, that is not the only way. We have lots of clever tricks to solve problems more efficiently than may seem possible at first glance—Fil-C is a bit of a modern marvel in that regard!
If you want to run everywhere where C does, you can't rely on that. I gave WASM as an example - that's a widely used target that just exposes a flat memory model where 0 is literally just a normal address and there's no way to trap it (unless they've released extensions I'm unaware of). Same deal with most microcontrollers as far as I'm aware of, although I don't do embedded.
I can't think of a way you'd implement null trapping efficiently on those platforms.
Words have very special meaning in ISO standards.
Fil-C is amazing and a prime example that undefined behavior means implementor freedom, and the implementor can choose to always trap on null pointer use. Sometimes the implementor freedom doesn't buy you much; for example why should it be UB to do
(const char*)NULL + 1
Dereferencing null is and should be UB but why is just calculating a pointer problematic? I just did some research and some old architectures would actually trap on creating an invalid address. So if we want C to support those machines, the standard can't define the behavior to do something other than what the hardware does.Btw, a pointer in C doesn't necessarily need to mean an address (invalid or not) in your underlying machine. C is a formally defined abstract language, not portable assembly.
That's an argument in favour of 'implementation defined behaviour'. Not 'undefined behaviour'.
IMO UB as "undefined but don't be crazy please" was the original meaning of the standard but people argue on that. It is a fact that compilers didn't exploit UB as strongly back then. However there is a good reason for this change: if you want formal semantics (which you do want, at least possibly) it is pretty much impossible to distinguish the two. If "undefined behavior" is undefined in the math sense, or in formal semantics of languages - the operation can reach any Abstract Machine state, then the fact that you cannot reason about anything follows immediately. The only dubious thing is time-travel, and this was indeed removed in the last version of the standard (and also for Rust now).
That's just compilers optimizing your code under the assumption that it correctly follows the rules of the language, which is the basis for any optimizing compiler and completely sane. The compiler isn't optmizing any null checks that actually correctly prevent all UB only ones that are faulty and never fully worked to begin with and as a result are indistinguishable form any other pointless code paths.
This is the bit that seems insane to me. Some of the behaviour of C/C++ compilers when they encounter undefined behaviour seems less like 'this is undefined, we'll just do something vaguely reasonable given the context/produce an error' and more like 'ahaha, the user has fallen into our trap, lets fuck them up'.
if (x > 0) {…}. But what if you entered the body when x <= 0?
if (false) {…}. But what if you execute the body?
These are “impossible”. What happens when the impossible occurs is “undefined”.
When the older standards said signed integer overflow for addition is undefined what they are actually saying is that the real definition of + is:
int +(int x, int y) {
assert(in_range(actual_math_add(x, y), signed_int_min, signed_int_max));
return machine_add(x, y);
}
So of course what happens when you get signed overflow is undefined; you should hit that assert and your program should explode and die. You should “never” get to the next instruction.But, in the interest of performance, “release mode” (which in this case is just any compilation) elides asserts since as a programmer you should not write code that asserts in much the same way that you should not write assert(false) in a normal code path that is supposed to run. Assertions are intended for “impossible” code paths and usually get compiled out in “release mode” though maybe your code is buggy and can actually hit them and then your program goes off the rails because it had a bug.
Put another way, if you did write assert(false) in a regular code path, would you find it unreasonable for the compiler to just delete the code after it? That is what undefined behavior is for.
For example, this code with an improper guard:
if (!p) puts("error");
printf("%d", *p);
Since the program dereferences p in line 2, and dereferencing null is UB, the compiler is allowed to assume p is never null, so it’s allowed to delete line 1, even though it would have executed before the point where UB would happen.Even worse, the compiler isn't just allowed to not do things you told it to do, it's also allowed to do anything too.
I was explaining why undefined behavior as a concept is a very sensible idea. Whether the expansive interpretation of the optimizations you are allowed to do when encountering the “impossible” are reasonable is a different question.
If UB meant what you described, it would be sensible, but it doesn't, and it's not.
UB is a formal tool for the optimizer.
Most people who argue about UB on HN have no clue wth it means and how it differs from Unspecified and Implementation-defined behaviours. The standard already explains expressions/statements and how they relate to sequence-points/sequenced-before/sequenced-after code points which is what is needed to understand the anomalous behaviour above.
Add in a introductory class in numerical analysis w.r.t. accuracy/precision/limits/rounding and the C/C++ programmer has enough knowledge to avoid problems in practice.
> if (false) {…}. But what if you execute the body?
> These are “impossible”. What happens when the impossible occurs is “undefined”.
By the way, that's exactly what happens with Spectre and its class of ghostly vulnerabilities! The processor speculatively executes these "impossible" paths, and discards the result once it detects they couldn't happen; but there are ways to "leak" information from that irreal world through side-effects like cache lines being discarded.
Rust's unsafe mode, for instance, has undefined behavior. A whole list of them, in fact.
There wasn't room for safe programming practices, and direct manipulation of the hardware was a design requirement.
It assumes that you know what you are doing.
There are also ports to the Zilog Z80, an architecture with similar limitations (UZI, FUZIX).
Also the decade predating C, already had high level systems languages in less powerful hardware, JOVIAL, ALGOL extensions, NEWP, PL/I subsets,...
> Surely the compiler can just check if each access is valid.
Well, sometimes it can, but sometimes it doesn't know how long the array is. What if the array is passed as a pointer?
> Maybe each array could be annotated with its size at runtime, and accesses could be checked at runtime too.
That works, but it adds runtime cost that may legitimately be too much for some applications, for example, a Gameboy game (set aside that many Gameboy games were written in assembly).
> Fine, so we'll make the programmer promise to ensure array accesses are always valid. Maybe they'll make a mistake sometimes, but what's the worst that could happen? Throwing your hands in the air and saying the compiler is allowed to do anything, that's just stupid.
Well, maybe it's stupid, but this is one thing that could happen if you accidentally write past the end of an array: https://www.youtube.com/watch?v=Vjm8P8utT5g. I'm sure neither the programmers nor compiler writers intended that.
Ultimately, the compiler can't guarantee any behavior if its assumptions are violated. The example may seem contrived, but it demonstrates that, given the right circumstances, the results of the logical contradiction are unbounded. This is a direct consequence of the "Principle of explosion": https://en.wikipedia.org/wiki/Principle_of_explosion. On second thought, maybe the runtime costs of array bounds checking are an acceptable trade-off after all.
int main() {
int x = 1;
a(x);
b(x);
return x;
}
Can I use constant propagation to replace the reads of x with the constant 1? This relies on the assumption that there is no UB, since otherwise our function calls could smash the stack and overwrite where x is stored on the stack.