Pointers Are Complicated, Or: What's in a Byte? (2018)
ralfj.de
ralfj.de
void *x = malloc(...);
free(x);
if(x); // undefined behavior
Note that this isn't about dereferencing x after free, which is understandably not valid. Rather the standard specifies that any use of the pointer's value itself is undefined after being used as an argument to free, even though syntactically free could not have altered that.This special behavior is also specifically applied to FILE* pointers after fclose() has been called on them.
If there is some historical reason / architecture that could explain this part of the specification I would be interested to hear the rationale, this has been present in mostly the same wording since C89.
void *x = malloc(...);
void *y = malloc(...);
assert(x != y); // standard guarantees this [1]
Yet it's fairly reasonable that: void *x = malloc(...);
free(x);
void *y = malloc(...); // malloc reused x's allocation here.
So, in effect, guaranteeing that the results of two mallocs can never alias each other, while allowing the implementation to reuse freed memory, requires semantically adjusting the value of a pointer to a unique, unaddressable value.[1] I think, but I'm not sure which versions of C/C++ added this guarantee
It's true that the result of "if (x == y)" would depend on coincidences lining up and you should not rely on either one. Calling any evaluation of x "undefined" seems much more extreme than that though.
Let's say you have two pointers that are sometimes unique and sometimes aliases. Maybe they mean semantically different things but they happen to be the same for some cases, and different in others. They always are on the heap. You want to clean them up when your function exits, freeing them both, or once if they are not unique.
free(p);
if (p != q)
free(q);
Believe it or not I have written something like this, although with integer file descriptors being closed rather than heap buffers freed. eg. Maybe some fds are passed for both reading or writing, or sometimes you have a unique fd for each, but it all needs to hit close(2).To exist within the standard I guess you need to do the comparison first:
if (p != q)
free(p);
free(q);
Edit: ah, but you just said "cross-object pointer comparisons are UB". I can't see a good reason for that either, but I do suppose it might make some architecture's non-linear pointer representation work better.It seems like they could have just said: malloc won't give you a pointer that overlaps with the storage of any live malloc'd object. Such a malloc is implementable without too much trouble. But instead, they gave a stronger guarantee--that all malloc'd pointers would be "unique". It would be unboundedly burdensome on the implementation to meet this property, so what do they do? Update the standard to offer the achievable guarantee? No! They add a new rule, ensuring that it's impossible to observe that the stronger guarantee is not met without doing something "illegal". Instead of getting their act together, they have elected to punish whistleblowers.
It certainly does not make sense for current implementations, and I find it difficult to imagine a pointer representation where it does make sense. Perhaps if reading the address itself involved some indirection.
Annex J.2 "The value of a pointer to an object whose lifetime has ended is used (6.2.4)."
6.2.4 "The value of a pointer becomes indeterminate when the object it points to reaches the end of its lifetime."
The only thing I can think of is that it allows the compiler to recycle the storage occupied by the pointer itself for something else after free() is called, but I don't see much value(!) in that either, given that if it could do variable lifetime analysis, it would be able to do it for pointers too.
One must keep in mind that the standard does not exist in a void, despite what the smartass language-lawyers unfortunately think/want. The standard committee simply did not decide to standardise things that they understood could vary between implementations. In fact, the standard has always said that "behaving during translation or program executed in a documented manner characteristic of the environment" is a possible outcome of UB.
I'm not aware of any architecture that does this, however, I think this is exactly how 80286 and later behave in protected mode, if malloc() allocated fresh segments.
He is talking about edge cases. I've been using C and C++ for over 20 years now, for both low level and high level stuff, and I never needed to know about those edge cases.
And I think you don't want to know about them, you don't need to know about them to write good code.
Nitpicking about edge cases tradeoffs made by compiler designers and language standards decades ago is like attacking under the belt, in my opinion.
It’s like a trolley problem in a way. The situation may be unlikely, but better have an answer for it ready instead of replying ‘how dare you ask such questions’.
Also, it is incredibility defensive to accuse the author of "attacking under the belt" when his observation wasn't aimed at any specific language.
An other way to say is that this is propaganda.
> Both parts of this statement are false, at least in languages with unsafe features like Rust or C:
The two languages share these features, which is the whole reason this is being written in the first place.
This is like calling code "ugly" and saying that's a feature.
Pointers are not unsafe.
https://stackoverflow.com/questions/61114026/does-stdptrwrit...
Also, he wrote a follow up article to deal just with uninitialized memory
https://www.ralfj.de/blog/2019/07/14/uninit.html
edit: the "other kind of byte" that is the padding byte is just uninitialized memory - so perhaps not another kind?
int x; // uninitialized
int y = x; // initialized? "frozen"?
int z = y ^ y; // zero?
I do not think this has been worked out yet–although I never get invited to these discussions, so perhaps I missed recent developments in this area. The other missing part is trap representations–I think Rust only has unspecified, but not indeterminate values. And finally, padding is different from uninitialized memory because it is quite difficult to actually initialize it–the standard lists some ways to zero it out (brace initializer lists) but it seems like pretty much any operation will move it back to being uninitialized. So that's a different than normal uninitialized memory for like static variables, since those have values that you can actually give them and they stick around rather than staying in this weird indeterminate state.It's UB to read from uninitialized data, and the thing about UB is that the compiler doesn't need to be specially consistent about it. It doesn't need to select the same y for y ^ y, it might transform this into 0 ^ 0 or 10 ^ 5 or select any other arbitrary value for each instance of y in this expression. So the value of z can't be relied upon. even int z = 0*y; can't be relied to produce z equal to 0!
The compiler might even delete the whole code after int y = x; because after this line the code is UB. That's why the compiler might remove null-checks (and all code associated with it) done after you dereference a pointer, because since it's UB to follow a null pointer (the only way the program could be possibly correct is if the pointer were not null!), all further checks become dead code.
Also, the UB story in Rust is still being developed, but regarding memory representation it's expected to be the same as C. That's because LLVM is quite happy in doing such atrocious but spec-compliant optimizations.
Also, now that LLVM supports freeze as an operation (see links below), Rust might surface freeze to the programmer as a primitive low level operation, in order to control the propagation of uninitialized-ness and avoid doing UB when reading variables. It would be something like
let y = mem::freeze(x);
https://internals.rust-lang.org/t/taming-undefined-behavior-...https://www.reddit.com/r/rust/comments/gnx7y0/update_to_llvm...
I’d always imagined that as parallelism rose, a new model of virtual machine would rise with it, with its own “assembly” (i.e. low level language closely tied to the machine), which would in turn be the target of a compiler of a new parallel-by-default high level language. Alas, this has not happened.
Chip design languages like VHDL might some day rise up from being slightly too close to the hardware to fill this role, but I certainly don’t know enough about any of these topics to have a qualified opinion.
Let's start with branch prediction. Modern hardware branch predictors are pretty phenomenal. They can predict a basic loop with 100% accuracy (i.e., they can not only predict that the backedge is normally taken, but can predict that the loop is about to stop). In the past decade, we've seen improvements in indirect branch predictor. The classic "branch taken/branch not taken" hint is fundamentally coarser and less precise than the hardware branch predictor: it cannot take any advantage of dynamic information. Replacing the hardware branch predictor with a compiler's static branch predictor is only going to make the situation worse.
Now you might be thinking of somehow encoding the branch predictor state in hardware. This runs into the issue of encoding microarchitecture in your hardware ISA. If you ever need to change microarchitecture, you face the dilemma of having to make a new, incompatible ISA, or emulating old microarchitectural details that no longer work well. MIPS got burned with this by delay slots. In general, the rule of thumb hardware designers use is "never expose your microarchitecture in your ISA."
A related topic is the issue of memory size. Memory is relatively expensive compared to ALUs: we can stamp out more ALUs than we know how to fill. We have hit the limit of physics in sizing of our register files and L1 caches (electrons move at finite speed!), and the technologies we use to make them work quickly are also prohibitively power-hungry. This leads to the memory hierarchy of progressively larger, but cheaper and more efficient memories. The rule of thumb for performance engineers is that you count cycles only if you fit in L1 cache; otherwise, you're counting cache misses, since those cache misses will dwarf any cycle count savings you're likely to scrounge up.
Now that there are languages that are higher level, people want to imagine that C is a low level language, when it actually isn't.
This broke me. I loved C++, I love pointers and shared memory, synchronization etc but optimisations finally creeped me out enough to switch to Java.
Not in the language itself, as the semantics are better defined, but on the overall performance behaviours, specially across different vendors and their respective implementations.