A reminder for all that have forgotten: UB is the one that can email your local council and submit a request to bulldoze the house you’re in. It is not a free core dump.
A reminder for all that have forgotten: UB is the one that can email your local council and submit a request to bulldoze the house you’re in. It is not a free core dump.
It is 100% well-defined behavior to dereference these pointers. It always segfaults, which as Jarred mentioned is a lot like a panic.
Rust evangelists need to be careful because in their zeal they have started to cause subtle errors in the general knowledge of how computers work in young people's minds. Ironically it's a form of memory corruption.
On Windows and Linux this is the first 4KiB so range 0x0000 up to 0x1000, unless large pages are on (then it's even more).
On macOS in x64 this is the entire 4GiB memory space, probably a method to help developers port their 32-bit software to x64. I don't know what the zero page size on ARM is.
If your microcontroller doesn't have this guarantee, you can't make use of this feature.
Remember this gem?
https://kristerw.blogspot.com/2017/09/why-undefined-behavior...
Once you trigger UB, all bets are off and your code could do anything. A segfault just means you spun the roulette wheel, bet it all on red, and got lucky your house wasn't bulldozed.
Zig also uses LLVM under the hood, right? So it's subject to these same semantics. An LLVM pointer value cannot legally contain arbitrary non-null non-pointer integers such as 0x2. That's a dead giveaway of UB. And I doubt the emitted Zig code safety-checks every pointer dereference for a value less than 0x1000 before performing the dereference.
Zig UB is not C UB. There is an entire language built on top of it. Just because something behaves a certain way in C, doesn't mean the same thing is true in Zig. Zig is no longer a code generator for C, it has switched to a self hosted compiler a while back. In fact, the language is rapidly progressing to the point where LLVM is a mere optional dependency.
I don't know the semantics around LLVM pointers. I don't see why 0x2 would be invalid, there are plenty of platforms programmed in C(++) that have a flat memory model. It would be quite painful to have a microcontroller where you can't send data to the output pin because LLVM decided that 2 is invalid (but 0 isn't). I've never seen LLVM complain about invalid dereferencing, though, it always ends up doing what the compiler tells it to do as far as I can tell.
Zig pointers will definitely cause UB but most Zig code shouldn't need them. Slices are actually bound checked and should probably be preferred in most cases of pointer arithmetic. Simple pointers can't be increased or decremented so you need to manually go through @intToPtr if you want to do real pointer arithmetic, which is quite unusable.
I haven't used Zig much so I don't know how many Zig semantics are copies of C semantics and how many are translated by the Zig frontend. However, "this is a bad/undefined thing in C so it must be a bad/undefined thing in Zig" is simply not true.
LLVM has rules about what is legal and what is not legal. If you follow the rules, you get well-defined behavior. It's the same thing in C. You could compile a safe language to C, and as long as you follow the rules of avoiding UB in C, everything is groovy.
Likewise, this is how Zig and other languages such as Rust use LLVM. They play by the rules, and get rewarded by well-defined behavior.
I will return the courtesy, with regards to my interpretation:
> An integer constant other than zero or a pointer value returned from a function not defined within LLVM may be associated with address ranges allocated through mechanisms other than those provided by LLVM. Such ranges shall not overlap with any ranges of addresses allocated by mechanisms provided by LLVM. [2]
[1]: https://llvm.org/docs/LangRef.html
[2]: https://llvm.org/docs/LangRef.html#pointer-aliasing-rules
- Any memory access must be done through a pointer value associated with an address range of the memory access, otherwise the behavior is undefined.
- A null pointer in the default address-space is associated with no address.
A null pointer (0x0) is associated with no address, therefore it has no address range. So if you do attempt a memory access (dereference), the behavior is undefined. QED. A naive translation to assembly would indeed segfault on a modern OS, but LLVM's optimizations are free to assume that code path is unreachable and do anything else.
Once the program is in this state, a bug of some kind is unavoidable. I don't take issue with that - what I take issue with is your claim that this behavior is well-defined, because it definitely is not. It would be equally valid for a null dereference to corrupt your program state or wipe your hard disk.
In Rust for example, derefencing a raw pointer is unsafe - because that pointer could have a value of 0x2 - which would result in undefined behavior according to LLVM.
tbh I'm surprised any of this is even up for debate. If you google "is segfault undefined behavior" you'll get 100 results telling you yes, yes it is.
I'm not sure how I can connect the dots any more clearly. Like gggggp said, it's baffling to see the creator of a popular language sweep the nasal demons under the rug and pretend that certain undefined behavior is guaranteed.
Calling such segfaults "safe" or "well-defined" is setting your users up for disappointment and CVEs, because a "well-defined" result is axiomatically impossible in the presence of undefined behavior. It's subtle, and if we were talking about a Java competitor maybe I could forgive the mistake. But if you're writing a low-level language it's important to understand how this stuff works. Ironically, he spread misinformation in the very post where he accused Rust evangelists of the same.
This thread is long dead and continuing the discussion seems futile, so I'll just leave it at that.
*excluding something silly like `raise(SIGSEGV)`
Do the docs actually define exactly which mechanisms external to LLVM count as allocating address ranges and which do not? It's possible that calling mmap and passing PROT_NONE does not count, for example.
I don't believe PROT_NONE suffices. The address needs to be accessible, not merely mapped. If reading through a pointer, the address must be readable. If writing through a pointer, the address must be writeable. This is why writing to a string constant is undefined behavior, even though reading would be fine.
Another issue is alignment. If you read from a `*const i32` with unaligned pointer value 0x2, the optimizer is free to assume that code path is unreachable and, you guessed it, bulldoze your house. If you get a segfault from reading an `i32` from address 0x2, you've already hit UB and spun the roulette wheel.
In theory the emitted code could check pointers for alignment and validity (in whatever platform-specific way) before accessing them, and simulate a segfault if not. Such checks would serve as optimization barriers in LLVM, and prevent these instances of UB. Of course Zig's current ReleaseSafe doesn't do this, and I think it would be silly if it did. But that's the only way you could accurately call segfaults "well-defined".
0x2 is a perfectly valid pointer value, it just happens to never be a good virtual memory address on modern systems where virtual memory is setup by the usual OSs, hence the fact that you can rely on it segfaulting.
Then I guess it could be a language guarantee if Zig only supports/targets those platforms. However, considering how low-level Zig is, I doubt that that is the case.
It isn't, in the general case. But JavaScript engines do some dark magic with pointer packing / NaN boxing as a performance optimization (most things in the VM are single words, passing around a single word is usually way cheaper than full unboxing), and I suspect bun in occasionally running into issues where it gets returned something from the JS engine which it thinks is a pointer but actually it's a packed, special value. This is a logic error, that turns into a weird memory issue at the abi boundary, not a memory safety issue.
Is that guaranteed by the language semantics, or could it possibly change at some point in the future? If it's the latter, then yes, it is very much Undefined Behavior, and not guaranteed to segfault before opening the door for potential exploits.
Not on every architecture, not in LLVM (even if well-defined on the underlying architecture), and not in C (even if well-defined in the underlying compiler backend).
TL;DR: Zig injects checks and aborts the program at runtime unless you specify that you wish to ignore the problem. This can be done explicitly within the code or by compiling under a build mode that ignores checks (unless specified manually).
Programs compiled as Debug and ReleaseSafe will terminate at runtime if UB is triggered. Compiling for ReleaseSmall and ReleaseFast will cause traditional C-style UB. If you care about your program doing what it's supposed to do, you use ReleaseSafe. Doing Release[Fast|Small] will do something similar to -O3 in other languages, which will often change behaviour.
Note, however, that you can compile your code under "just allow UB and see what happens" mode but still benefit from checked UB by setting @setRuntimeSafety(true); this will introduce the assertions despite the unsafe build modes you may specify.
It's like introducing a C++ compiler flag* telling the compiler "ignore exceptions and just continue". You know you're in for a bad time the moment you specify it, but it makes your program blazingly fast because it greatly reduces the amount of code to generate/checks to execute.
The main advantage of checked UB is that well-tested code can make use of the unchecked nature of these features for speed without having length check code blocks that need to be wrapped in debug #ifdefs or similar. Assuming you don't run test builds with checks enabled (and why wouldn't you) you'd catch these problems in your build pipeline.
This is different from the normal way of working with C and friends, where UB remains in debug/-O1 builds but just acts a little differently. Some compilers will insert breakpoints, others will ignore the problem like in release mode, nobody knows what will happen and your compiler can't detect this problem for you.
* note that -fno-exceptions exists, but that aborts the program rather than let it continue.