Undefined behavior is often a thing for C programmers (2011)
blog.llvm.org
blog.llvm.org
PS: some pragmatic cleanup of UB and IB in the C standard would be welcome though.
As soon as your first pass decides to rely on the presence of the dereference as a reason for removing the null check, the dereference needs to be flagged as critical, such that it cannot be removed by a subsequent pass.
Better still, if the compiler does try to remove something that was relied on by an earlier optimization step, alert the user that something is probably wrong with the code.
Also, this would prevent one of the main benefits of function inlining. Once you have inlined a function, the values of some of the parameters may be known at compile time, resulting in lots of those if(0) parameters. But that is something that was relied on in previous passes, even though it is known to be a static value in the current pass.
UB is a feature, not a bug. So you don't really "clean" it up. Cleaning up of UB means a lot of overhead for programs.
for (i = 0; i < *len; i++)
array[i] = 0;
without having to check on each iteration whether len secretly points inside the array and got modified by the loop. It would be better if this were not a heuristic and there were clear compile-time information saying for sure that this optimization is safe.In case of similar but where it’s float array[N] etc then compiler can assume that they do not alias.
One of the reasons why Fortran was faster for decades was the default aliasing rules, restrict finally got us out of that trap. If we'd always assume everything alias plenty of code would become way slower.
Also C has well defined ways to access memory via different types. Make an union type and access via that. It makes it explicit.
Since then they specified that unions are the way to go.
It's quite common in existing real-world C code to rely on the ability to fill in a variable via a pointer of one type and then access it via a pointer of another type, and you expect the language to properly return the data you wrote to that pointer. The C committee determined that, just about always, that happens when one of the pointers is a char pointer (e.g., you call read() or write() on a struct), and they said that char pointers are allowed to alias by default but others can't. That's their guess, which is mostly accurate. There are much nicer answers here if you're not obligated to be compatible with existing C code - but in that case, especially today, I'd go with one of several other languages that can accurately express intention at the source code level instead of making backwards-incompatible changes to C. :)
Some examples:
1. When C came out, you couldn't assume a two's complement machine. So some architectures would yield different results on signed over/underflow than others, and some architectures might trap on overflow. Hence, signed overflow is undefined in C.
2. Did you know the Lisp Machine had a C compiler? It did. Because pointers on the Lisp machine were tagged and bounds-checked to ensure memory safety, this compiler implemented C pointers with fat pointers, containing a reference to an object and an offset within that object. Most machines you are familiar with allow indirecting through arbitrary integers; the Lisp Machine is an example of one that does not, trapping on invalid pointer accesses. Having C and C programs available in both environments mean that indirecting through a pointer that does not point to or inside some pre-existing static or dynamically allocated object is undefined.
3. Operating systems with virtual memory can arrange things such that indirecting through 0 (NULL) will raise a segmentation violation in user code, guaranteeing a pointer value that is always invalid -- but inside the kernel you may want to access memory location 0. So accessing null pointers is undefined.
By leaving these undefined, you are not committing implementors to adapt the abstract machine to the underlying hardware in cases where hardware can widely differ, allowing for simple implementations that yield fast code on all architectures they compile for yet still allowing for a useful subset of C code to be 100% portable.
Or just does what C does, make it UB and the programmers will be aware of implementation of overflows and won't suffer because of added checks. They will do the checks themselves if they want and handle their numbers with care
Imo the answer always boils down to performance. Java has no UB yet it works on billion devices (citation needed). It is doable, just costly.
MIPS compilers for C emit "addu" instructions for signed integer additions rather than using "add", to avoid overflow traps.
Using unsigned arithmetic for signed is not just so that C programs relying on the behavior do not blow up. When overflow behavior is reversible wraparound, the compiler more freedoms in rearranging a calculation. Suppose the programmer has written some a + b + c, and ensured that there is no overflow (when it's added left to right, as the abstract semantics says). If the addition mechanism has reversible overflow, the compiler can rearrange it to do the additions in any order. The result will be the same, correct one that (a + b) + c would produce.
But the idea that two's complement itself "wraps around" is mostly a myth. Two's complement has a fixed range of representation, and a result of adding two elements from that representation can be out of that range. In that case, the truncated value whose bit pattern corresponds to the unsigned arithmetic is not the correct result.
Even if such exotic hardware is still around, the only thing that would change is that C compilers supporting this hardware couldn't call themselves "ISO Compliant" anymore, the users of those compilers wouldn't lose anything.
Nobody during C89 formulation realized that it would end up “a license for the compiler to undertake aggressive optimizations that are completely legal by the committee's rules, but make hash of apparently safe programs” that “negates every brave promise X3J11 ever made about codifying existing practices, preserving the existing body of code, and keeping (dare I say it?) `the spirit of C.'” If they had, the response would have been that “[UB] must go. This is non-negotiable.”
I’m not quite sure what the compiler is supposed to do in that case. Insert nullchecks by itself to avoid crashes?
How can this be: C has a standard, when many other languages do not. Thus C is "more defined" than other languages. Yet somehow it ends up being less defined.
If C had no standard, it would be defined by its implementation, just like many other languages. But because it has a standard, it has undefined behavior.
The real trick is that without a standard, implementations have to be bug-compatible. They're too defined! The purpose of a standard is to define them less.
But for most intents and purposes Python is defined by its documentation, and it's fine.
Having said that UB in C and C++ means behavior that the standard doesn't define. If you remove the standard then UB either loses meaning or it makes all programs undefined. It definitely doesn't make all programs defined.
If C had a judge, like the legal system, he could deny attempts to use fuzziness in the standard (intended to prevent specifying bugs) to introduce bugs. Unfortunately, C doesn't have a judge with common sense, it has people who think the standard is code and anything that is undefined really is a license to do whatever you want.
I guess this problem was inevitable as soon as the words "nasal demons" were put to keyboard. Now the only thing that can save us is a reference implementation.
You'd hate a compiler where this has silly overhead:
for(int i=0; i < n; i++) arr[i];
That's because 64-bit addressing modes don't overflow the same way as signed 32-bit ints. But signed overflow being UB gives compilers license to ignore the difference and generate code that makes sense.The real problem is that C is too bare-bones and lax to be able to contain or avoid UB in day to day programs. For the aforementioned overflow gotcha, there's no way to enforce use of size_t for indexing. There's no `checked_mul(a,b)` to easily avoid signed UB. You have to make the check yourself, and if you're not careful, that check will be UB too. malloc() makes you multiply numbers often, exactly where it's dangerous. And on top of that implicit integer conversions can make unsigned arithmetic surprisingly signed.
This article [0] goes into a bit more detail.
[0]: https://gist.github.com/rygorous/e0f055bfb74e3d5f0af20690759...
It's trivial to prove in that particular example because it's assumed that count is also an int (both i and count are in 32-bit registers).
> It also claims that the compiler cannot prove this in more complex examples (but gives no explicit examples).
The example doesn't even have to be that complex; you just need sizeof(n) > sizeof(i) and i to be signed (godbolt example at [0]). In that example, if the compiler can assume signed overflow is UB, then you get fairly sensible code generation. If the compiler cannot make that assumption (add the -fwrapv flag), then you get that movsxd instruction in the middle of the loop body. If you use -O2 instead of -O1, then -fwrapv prevents loop unrolling.
> What I'm saying is if the compiler cannot prove a lack of overflow in this way, that should be a warning.
I'm curious how frequently this warning might fire in real-life codebases, or when it does fire how frequently it's something the programmer can actually fix. Clang does have -Rpass-missed to emit some diagnostics when an optimization cannot be performed, but it doesn't emit anything in this particular case (though I'm not sure whether it's supposed to).
There's surely cases where it would be hard to suggest a "equally efficient without UB" replacement (eg when it wants to fuse two i64 into one), but it feels like the current position is too far on the axis of "provided with suspicious code, the compiler should just make that garbage run as fast as possible and call it a day"
Compilers incorporate inferences that can be made by assuming UB doesn't occur into their optimization because it improves code generation of correct programs.
I wouldn't say it's a minuscule factor in my daily work, as OP put it, but it's not something to agonize about if you stay mindful of avoiding common undefined behavior pitfalls. It would be very helpful if compilers were better about warning you when you're venturing into UB, but that's a different topic.
1: You can even get the shirt: https://teespring.com/shop/undefined-behavior-shirt
Many people say this but nobody "literally" believes it.
It's a way of expressing the righteous dominance of one person or group's attitude towards reasonableness over another's. In a way that forestalls argument.
"Literally anything" could literally mean global thermonuclear war.
If you get Godbolt to run the program and show the output, you'll see it change with GCC version.
Obviously what is "sensible" and what is "surprising" is subjective.
int foo(FILE *f) {
int val;
fread(&val, sizeof(val), 1, f);
return val;
}
Closer reading of the spec says something about it being fine to cast, say, an int* to a float* and then cast back before dereferencing it. In this example, we're OK, because we don't reference via the void* . But the fread implementation must do, right?I'm left feeling like it is impossible to implement fread() in C and that the standard library's API is no longer considered a good example of how to construct an API.
I guess the original motivation for it being UB to cast is to do with alignment constraints on some hardware. For example, a machine might be able to read an int8_t from any alignment but might require 4-byte alignment for int32_t. If so, then casting a non-aligned int8_t* to int32_t* and deferencing it would indeed fail on that hardware.
But some would argue that it is not right that GCC 7 breaks such code on a machine where that problem doesn't exist.
People should write these bits in assembly (as little as possible, but doing the casting and things), and then keep their C code well-defined using these low-level library routines.
The key difference is that the alignment (usually) cannot be determined at compile time, so the compiler would still have to generate code that worked if the alignment was correct at run time. I think this is what many people either expect or would like the compiler to do.
With that change to the spec, I could continue to write my code in portable C, and I'd only get an abort on the target with this alignment constraint.
What if on some other machine it doesn't cause the program to abort but instead returns a broken result? How do you cover that in the spec?
> But some would argue that it is not right that GCC 7 breaks such code on a machine where that problem doesn't exist.
Right, so let's write these low-level things in C and not worry about the case of 'where the problem doesn't exist'.
OK, how about, "dereferencing a pointer that isn't correctly aligned doesn't work on some hardware". I see that this is a hard area for the spec to do well. But the current situation isn't viable. I can't tell whether I can safely use fread for heavens sake.
I feel like dereferencing an incorrectly aligned pointer could be a runtime error, just like divide by zero or accessing memory that doesn't exist. Maybe that would be an improvement from the current status. Although I'm not sure the spec _could_ be changed like this at this point in time.
By checking the alignment on every access? That'd be far too expensive if the machine you're on doesn't trap.
Editing here because it looks like we've hit the max reply depth. This wouldn't be like UB in the same way that code that could generate an out-of-bounds memory accesses is not UB. The critical thing is that the compiler should assume that the hardware error does not occur.
Yeah, but good luck selling that... for the reasons you've identified it'd be a nightmare to run C on it.
You can cast void* to something else and it will work, but you can't expect to get a random void* and presume it has some valid well defined data for you to read. You must know where it came from and what's been done to it.
As a rule of thumb if you just write dumb bytes it's char*. If it's anything else you must know what you're doing and there are ways to perform type punning that's within the spec.
Can you expand on that? I don't understand.
Edit: I see the answer in one of your other comments:
> If len and array are both of the same type then the compiler must assume they can alias (or one of them is char*). One has to explicitly state with restrict that they cannot alias.
I'd happily take that hit. I can put restrict in the hot paths of my code. I'm tempted to say I want restrict to be the default, but I haven't thought that through.
I know. I think the reason that this an area that causes lots of debate is because the majority of C programmers from 20 years ago would not have thought the memcpy or unions approach was a better idea than casting and dereferencing void* . The casting approach used to "just work" unless you violated the alignment requirements on your platform, and that was a problem 20 year-ago C programmers were happy to take on.
For next time this discussion comes up, I'll try to think of a good example of when the void* approach yields easier to maintain code :-)
BTW, you can also use -fno-strict-aliasing to make the C compiler sort-of work like it did 20 years ago. I believe this is what the Linux Kernel does but I couldn't be certain that it is for exactly this reason.
If you're certain there is no violation (such as with malloc), then go ahead and do it. -fno-strict-aliasing is unrelated to this issue. It removes the strict aliasing assumption from the compiler (another source of potential issues if not understood), but void* (and char*) are explicitly allowed to alias to anything regardless.
I modified the code to ensure that the data is 8-byte aligned (I think I've done this correctly). It still goes "wrong". https://godbolt.org/z/Pnxnb8
In either case, adding -fno-strict-aliasing fixes it.
uint64_t *a = malloc(8);
What you're running into in the context of your example, is breaking strict aliasing.You're starting with uint32_t pointers, which then go through void pointers and end up being accessed through pointers to uint64_t. This breaks strict aliasing since void* can not be used to bypass it.
If you had started with uint64_t, you could go through void* and back to uint64_t without any issues.
And in the case of malloc, I don't know what type the data was "originally" before it returned it to me. I guess I have to treat malloc as a special case where I can cast its void* to anything safely. That would seem morally wrong.
You need to have a model of how the compiler/optimizer works. In your example, it's important to know what objects you're working with and what the effective types of these objects are. Focus on the objects, not on the pointer types that you use to access them. Normally, the effective type of an object is its declared type. Allocated objects however, have no declared type according to the standard [1]. They get their effective type on first access:
"If a value is stored into an object having no declared type through an lvalue having a type that is not a character type, then the type of the lvalue becomes the effective type of the object for that access and for subsequent accesses that do not modify the stored value."
This means that you don't care how malloc works internally (unless you're writing one), and casting/accessing its result - which is guaranteed to be properly aligned - does not break strict aliasing.
But in your example, the strict aliasing violations are crystal clear. You declare uint32_t objects, therefore the compiler knows and can track the effective types of these objects. Later, you're trying to access the values of these objects through different types which is when the strict aliasing violation occurs.
> So is the rule that it is OK to cast a pointer back to its "original" type and dereference it?
Since you're accessing the same object and there's no mismatch between its effective type and the type you're using, you're ok.
> If so, I feel this rule breaks the abstraction mechanisms C provides. eg now when I'm implementing that xor512b() function, I need to know about the invisible "original" type of the arguments.
void* is not really an abstraction and should not be used as such (or to avoid thinking about designing a proper interface). You can make use of the type system by declaring proper types and being consistent, or you can use char* as a byte-level interface. Usually, when you see void* in an interface, it's clear what the implications are and how you should proceed with it. If it's not, you're looking at bad design.
[1] http://www.iso-9899.info/n1570.html Footnote 87) Allocated objects have no declared type.
The fread implementation sees dest as a void pointer and will access it going through char*, which is explicitly allowed.
For the most part, avoiding UB is common sense when you grasp the basic abstract C memory model - types of storage and their lifetime, avoiding out of bounds array access, keeping track of the lifetime of all allocations, etc.
Do all C programmers know what strict aliasing means?
How about pointer provenance?
Have they memorized integer promotion rules?
These are all potential sources of undefined behavior that can lead to catastrophic results.
I bet most experienced C programmers with decades under their belt would fail one or more of these tests.
- More applications are talking to the internet and being exposed to malicious input. There's also more money to be made in finding exploits, and more people and governments looking. At the same time, fuzzing techniques are getting dramatically better.
- Applications are more likely to run across different architectures (mobile) and different compiler versions (faster pace of compiler development), leading to more opportunities for "something weird" to happen. Multithreading is also more common.
- Memory-safe languages have taken a larger market share, and the popular attitude towards memory safety issues is changing from "bugs are a fact of life" to "this is a drawback of specific languages".
https://raphlinus.github.io/programming/rust/2018/08/17/unde...
[0] https://blog.llvm.org/posts/2011-05-13-what-every-c-programm...
[1] https://blog.llvm.org/posts/2011-05-14-what-every-c-programm...
[2] https://blog.llvm.org/posts/2011-05-21-what-every-c-programm...
In my years of working with C its always the 'look how smart I am code' causes issues.
Keep it clean and simple. Avoid complexity and fluff features of C. And you will not going to have deal with exotic bugs.
Why would you not actively avoid undefined behaviour and instead take risk.
Zeroing out variables, checking pointers coming from external modules etc. Not rallying on a compiler to resolve fancy stuff.
Like the example:
void contains_null_check(int P) { int dead = P; if (P == 0) return; P = 4; }
Leave declarations uninitialized or zeroed Check input. Work with data.
void contains_null_check(int P) { int dead; if (P == 0) return; int dead = P //work P = 4; }
This leaves less space for error/interpretation and IMHO it keeps code cleaner.
Any fancy optimizations should be used only for critical path and only when absolutely necessary.
Quoting from the abstract of that talk: "... preserving source-level dependencies through to the generated CPU instructions is achieved through a delicate balance of volatile casts, magic compiler flags and sheer luck. Wouldn't it be nice if we could do better?"
In a past life, I depended on signed overflow to detect out of bounds memory allocations. That kind of stuff gets compiled away these days, you need to be cleverer to detect when the user is deliberately trying to invoke UB. And because the compiler doesn't believe in UB, there's the risk it's going to remove your detection logic.
IIRC C++ just decided that signed numbers are two's complement and that's that. C could really use similar thing.
Signed integer overflow is still undefined in C++.