That way, any programs that are made to work on the pathological implementation are largely guaranteed to rely only on standardized behavior, and work on any other standards compliant implementation.
That way, any programs that are made to work on the pathological implementation are largely guaranteed to rely only on standardized behavior, and work on any other standards compliant implementation.
A long time ago, when encountering "undefined" behavior, it would try to launch Nethack[1]. (It is valid behaviour, according to the spec.)
Seriously, ask around. There was a multi-month fat-chewing about gcc's pathological interpretation of "undefined" earlier this year on the Cryptography list, and tends to come up anywhere C programmers with an interest in security drink.
gcc 1.17 would invoke nethack (or one of several other similar games, if available) if it saw an unrecogized #pragma directive.
According to the C standard, the behavior of a #pragma not followed by STDC is implementation-defined, not undefined -- which means that an implementation is required to document its behavior. (I presume that gcc did so.)
gcc did not play launch Nethack in response to undefined behavior in general, or in response to anything other than an unrecognized #pragma directive.
So sub-optimal but predictable and safe code given that you change the preference bit to safety over speed.
A safe compiler will have to add checks to guard against undefined behaviour happening at runtime.
I had the same opinion at first and would agree if UB was mostly possible to avoid in meaningful programs, but it isn't. The main casualty would be the reason why this is being discussed at all: compilers need to make very specific assumptions on UB to enable some optimizations and if UB goes away, so do the optimizations. So we get the performance of the "boring" compiler anyway, just with more effort on behalf of the programmer.
DJB is (as always, I guess) right, C simply isn't really suited to optimizations based on UB. What'd be the big deal if compiler developers just stopped inflicting these on C programmers and focused on optimizing compilers for better-defined languages (with range types, specified overflow behavior etc.) instead?
This is non-trivial. These definitions come with tradeoffs that fundamentally alter the language. However, this is basically what rust is.
You're acting like it's unambiguously bad, but that's not true. For example, I might create a macro DEREF_AND_FREE(x) which expands to if(x)free(*x). Often times, I'll lazily use that in places where I know x isn't null. It's more readable and maintainble than splitting the macro up into two separate macros, DEREF_AND_FREE_IF_NONNULL and DEREF_AND_FREE_I_KNOW_ITS_NONNULL. These UB-based optimizations which are "inflicted" on me mean that I don't take the performance hit for my laziness.
$ cat test.c
int main() {
return *(int*)0;
}
$ gcc test.c -fsanitize=undefined -o test
$ ./test
test.c:2:9: runtime error: load of null pointer of type 'int'
Segmentation fault
I'm pretty sure I've seen similar tool based on clang as well.And there is valgrind.
Yes, that's hardware, but a nice place where hardware ain't C.
I'm just sayin'...
The latter is nothing unusual, x86 has it as well.
So if the board designer wired everything up to map RAM at address 0, then yes, you'd be able to dereference a pointer to address 0.
Completely beside the point, though. Dereferencing NULL is always undefined behaviour in C.
Maybe everyone who wrote real-mode systems software either ignored that bit of the C standard, or they fudged around it (IIRC the IVT is an array and the first entry doesn't matter, so you can start a few bytes past zero), or they used assembly to access it.
And of course dereferencing 0 isn't impossible in practice on arches which support this. Compilers for such machines usually can be coerced to generate code which accesses 0, you just have to live with the fact that your code isn't considered "valid, portable C" anymore.
long zero = 0;
struct IVT * ivt_ptr = (struct IVT *)zero;
and struct IVT * ivt_ptr = 0;
/* or (struct IVT *)0 if you like that better stylistically */
The former gives you a pointer whose bits are all 0. The latter gives you a null pointer. On systems where 0 is a valid memory address the compiler should pick some other address that is not valid.Similarly, these two are not necessarily the same:
if ( ivt_ptr == 0 ) ...
and if ( ivt_ptr == (struct IVT *)zero ) ...
/* or if ((long)ivt_ptr == zero ) ... */
0 only represents the null pointer in C when it is the constant 0.See this for more information: http://c-faq.com/null/
Not necessarily. Null dereference is undefined behavior so it may as well access some valid memory.
Even on platforms where 0 is a valid address, practical compilers tend to use 0 as null constant because it simplifies implementation (easy to check for null, matches the C syntax without WTFs).
This comes at the cost of making some small part of memory unusable to portable C code (nobody will believe you when you return 0 from malloc if 0 happens to be considered null), but that's fine because such memory is typically used for machine-specific, nonportable stuff anyway.
It's also perfectly fine on x86, in kernel memory page 0 is often mapped, the OS is the one mapping page 0 so that userland dereferencing it faults.
In fact, Linux had several privilege escalation bugs which involved putting something at 0 and executing a buggy syscall which loaded this thing due to NULL dereference and believed it's some legit internal kernel data.
* Well, this is implemented by the default system toolchain. You'r particular toolchain can opt out of this behavior if they want to.
Or it could document and provide a well defined behavior of derefrencing a NULL pointer.
(e.g. gcc provides -fdelete-null-pointer-checks to control this)
as(1).
1. you're not required to use C
2. an extension to the C standard can decide to define that behaviour
(void*)(intptr_t)0
can in principle be more involved than a noop.And an optimizing compiler is free to transform its generated code based on the assumption that the code's behavior is defined.
An optimizing compiler is free to transform its generated code based on the assumption that undefined behaviour never occurs.
""" As we say, the volume is almost all on the surface. Even in 3 dimensions the unit sphere has 7/8-ths of its volume within 1/2 of the surface. In n-dimensions there is 1–(1/2^n) within 1/2 of the radius from the surface.
This has importance in design; it means almost surely the optimal design will be on the surface and will not be inside as you might think from taking the calculus and doing optimizations in that course. The calculus methods are usually inappropriate for finding the optimum in high dimensional spaces. This is not strange at all; generally speaking the best design is pushing one or more of the parameters to their extreme—obviously you are on the surface of the feasible region of design! """
The issue is design axis's are entangled. And also figures of merit are nonlinear. Twice as whatever doesn't mean twice as good.
AVR Processors don't have bus faults. Some machines I've worked on have RAM based at address zero.
warning: indirection of non-volatile null pointer will be deleted, not trap [-Wnull-dereference] return (int)0; ^~~~~~~~ note: consider using __builtin_trap() or qualifying pointer with 'volatile'
That's a sweet little optimization :)
Probably the only reason they don't abort compilation at this point is that someone complained when it broke his tricky little macro which sometimes generates unreachable null dereference. Or something like that.
int.c:3:12: warning: indirection of non-volatile null pointer will be deleted,
not trap [-Wnull-dereference]
Much better, scan-build: int.c:3:12: warning: Dereference of null pointer
return *(int*)0;Signed integer overflow is the most common one, since it means you can hoist 32bit ints into 64bit ints so they fit in one register or similar and save hitting the stack.
I was about to say that "this would break a ton of other projects"—but actually, it seems like it'd be fine as long as it was an opt-in -W switch.
Let's say that, on the pathological compiler, it's documented that sizeof(int) is 4, and chars are signed if the number of seconds in the current minute is even, otherwise sizeof(int) is 5 and chars are unsigned. If your code compiles and works with either setup, it's probably correct and portable.
It's not really feasible for programmers to avoid implementation-defined things like sizeof(int) varying across implementations, not is it feasible for compilers to detect assumptions made about the implementation's definition in the program. The existence of a pathological compiler would make it far easier to write programs that work across all implementations.
It's true that working with unpredictable systems is a nightmare, but that's precisely the nature of undefined and implementation-defined behavior in C. The existence of a pathological compiler would merely allow developers to work through these issues on a single machine, which would be an improvement on needing to obtain a variety of compilers and platforms test with to find these issues.
Or you want the de facto widely understood behavior to happen which has happened on every compiler you've used in 30 years: at least every compiler that was for two's complement matchines. Or every compiler on which pointers to different data types were of the same size. And so on.
You also don't want to be burned by code that is relying on a common extension, when that is ported.
There is a word for the compiler approach of "we're going to optimize this based on the assumption that the program is not relying on a common extension that is technically UB".
That word is: malpractice.
For instance, it is not unusual for C code to be targetting machines in which all objects are in the same address space, such that pointers are internally like binary numbers, allowing pointers to different objects to be compared for inequality: ptr1 < ptr2 /* does ptr1 point to a lower address than ptr2 in the one big address space? /
Programmers expect this work consistently with the address structure of the machine. They don't want the expression to be somehow wrongly optimized based on the assumption that ptr1 and ptr2 must point to the same object.
Implementations must honor requirements beyond those of ISO C. Such as, for instance, behaviors that those same compilers* used to define historically in their past revisions.
If GCC has behaved a certain way for 25 years, and some code has come to depend on that, and then that is suddenly taken away, such that entire GNU/Linux distros continue to compile, but break in random places because of that detail, the fault for the regression lies 105% with the doofus who made the GCC commit. Even if that behavior is not defined by ISO C.
Member function pointer size varies widely between ABIs.
I've never seen C++ code which depended on the internals of pointers-to-member, or assumed they had the same size as other kinds of pointers.
A pointer to member is actually a complicated offset into an offset into some structure associated with an object's class. This offset can be "dereferenced" with respect to objects of different types in the same hierarchy, taking into account multiple inheritance and virtual bases, etc. A pointer to a Base::Foo function, can be applied against a Derived object, such that it resolves to the correct code, even though Base is the third base of Derived, and so Foo is in a totally different vtable position in Derived compared to Foo. (The pointer cannot be a blind integer offset.)
In C, we would make a struct containing function pointers, and a simple "pointer to member" would just be an integer "offsetof(our_struct_type, some_function)". Then worry about it later if someone wanted multiple inheritance. :)
For instance including a platform header like #include <fortran.h> is undefined behavior. ISO C has some requirements there: if the header is found, then the directive is replaced with the preprocessing tokens from reading that header (and so if there are syntax errors there, they have to be diagnosed, and so on). But ISO C specifies no requirements as to what <fortran.h> contains, or whether it exists. On one implementation, it might cause the rest of the translation unit to be treated as a Fortran program. On an other implementation, it might bring in some declarations related to a Fortran interoperability. On yet another, translation might stop with "header not found: fortran.h".
UB is simply "behavior upon use of a nonportable or erroneous construct, of erroneous data, or of indeterminately-valued objects for which [ISO C] imposes no requirements".
Nonportable is not necessarily erroneous!
The former is similar to the latter, only it must be documented. Unspecified behaviors imply some choice from among alternatives which need not be documented. (Like the order in which the argument expressions of a function are evaluated.)
The "boring" compiler doesn't exist yet, and would need to be implemented on every platform that currently has a standard compiler to be as useful.
These days, valgrind is a much better option for doing this.
Why not just fuzz inputs if this is your goal?
Think of, say, integer overflow. The problem is that typical compilers have a consistent (but non-standard) implementation, so it can't be fuzzed without some compiler help.