C was not created as an abstract machine
utcc.utoronto.ca
utcc.utoronto.ca
Because while all CPUs running C code have defined semantics for any given construct, not all CPUs have the same defined semantics.
Making C adhere to one CPU architecture for the sake of convenience would make implementations for other CPU architectures decidedly inconvenient. Rather than dereferencing an invalid pointer doing "whatever this CPU does", suddenly all C implementations now have to do whatever, for example, a VAX did.
It would mean that C code running on simple CPUs without memory protection would have to look out for invalid pointer derefs (how?) and simulate access violations if one were mandated by the spec. Or, if the spec said that the address space should be treated as a flat memory area all of which is accessible by any pointer, modern implementations would have to simulate that.
Dereferencing an invalid pointer doesn't do "whatever this CPU does," it does completely unpredictable things. I think that's part of the problem (to the extent it's a problem)
I recall IBM’s AIX used to put a zeroed-out block at address zero, so this code was guaranteed to always return 0:
*(int *)NULL
The reason was that it allowed their optimizing compiler to speculatively dereference a pointer before executing a surrounding if condition.It’s not up to the platform. If it looks like a code path might dereference a null pointer, the compiler can and will make wild and bizarre optimizations in the surrounding code.
Yes, some compilers will make "wild and bizarre" optimisations (they are neither) based on the assumption that any pointer which is dereferenced is not NULL - because if it were, the code could have done the exact same thing anyway, because there are no constraints in that case.
But compilers don't have to do that. There are no constraints. A compiler can emit code to do anything at all in the event of a NULL pointer being dereferenced, including emitting the machine code you'd naively expect - which may trap on some CPUs, or always return 0 on reads and swallow writes on other CPUs. It may also emit code that wipes your hard disk (and still counts as a conforming implementation). Or make demons fly out of your nose.
A compiler vendor may choose to guarantee a specific behaviour for some invalid constructs. They are allowed to do this because there are no constraints on what they can do. It's just that most don't, because 99% of people comparing compilers care more about benchmarks than they do about what happens with invalid code.
So in some ways, it is up to the compiler/platform/CPU what actually happens with code that has invalid behaviour. But as a C author, you can't assume any particular behaviour, because it might change from one CPU to another. Or from one OS to another. Or from one compiler to another. Or from one version of one compiler to another.
It basically means that the whole portability argument that is supposed to be in favor of C is just wrong because every compiler and platform actually brings its own C dialect and it is just sheer coincidence that it works at all.
Not sure about the C code bases you've been spending time in, but in the ones I've looked at, the overwhelming majority of the code has been well-defined by either the C standard or the implementation. In the cases where it hasn't, the vast majority of those cases were genuine bugs that needed fixing, and the fix made the code well-defined.
The number of codebases I've seen that actually relied on unspecified behaviour, or on whatever the current compiler/OS/CPU happened to do with undefined behaviour, is miniscule. (Or an entry in an obfuscated/underhanded coding contest.)
Whether things should be UB or implementation defined requires a good value judgement. Per se there is not a real need of UB, (assuming that the behaviour of a machine can be specified), but on the other hand there isn't much value in specifying a situation where all control is lost to the point where implementation details leak (runtime structures are overwritten etc.).
Since there is (I assume?) an expectation for "implementation-defined" to have an actual definition, pedantically an implementation leak would require all implementation details to be codified, and thus set in stone.
By making things undefined we can avoid a large amount of complex code to detect a situation that shouldn't happen.
Since a compiler cannot predict in all cases in advance whether a pointer is invalid it should be required to compile a pointer dereference the same way in all cases, i.e. emit the same code regardless of whether it has determined the pointer is invalid or not. Otherwise it effectively emits random code depending on how the optimizer is feeling.
Non-deterministic behavior for the same operation on the same values in the same program is not very helpful behavior, whenever and to the degree it can be avoided.
Whether this is a good or bad thing is of course a legitimate (and good!) question, but for writing C today that's how the language is specced and a reality the programmer needs to take care to avoid, just like a bunch of other things that C leaves to the programmer like remembering to clean up allocated resources when they're no longer needed.
Any time your are dealing with data from real, physical sensors or third-party APIs, this is an impossible guarantee to give - they could literslly break.
C is an absolute minefield of undefined behavior, but let's be accurate about the things it does wrong.
C23 somewhat improves the situation with <stdckdint.h>, a standardized version of GCC’s __builtin_add_overflow and friends. That has ergonomics issues too due to its verbosity, but at least it’s hard to screw up.
if (x < 0 || y < 0) return -1;
return x * y;
Will not overflow. And if you try to check for overflow with: if (x < 0 || y < 0) return -1;
if (x * y < 0) return -2;
return x*y;
Then yes, the compiler is within spec to remove your check because the only situation in which you could hit that check would be after signed integer overflow, which it is allowed to assume won't happen.One way to implement this check in GCC where the compiler will respect it would be:
if (x < 0 || y < 0) return -1;
int z;
if (__builtin_smul_overflow(x, y, &z)) return -2;
return z;Sure, but it didn't have to be like this. They could have said it is unspecified wihout allowing the compiler to assume it doesn't happen.
Would C have been better if the spec was different?
Even something like register allocation requires knowledge of what pointers point to.
It would be more helpful for the compiler to assume that it does not know what will happen. This is how C actually worked for many years.
> a C program executing undefined behaviour is _invalid C_.
That was not formerly the case, and it is not always helpful to redefine C in this way. Sometimes you really are not trying to write portable code, and you really do want the behavior you know that the target machine will give you, even if the C spec doesn't require it.
If it is really necessary to generate random code when some anomalous situation is encountered, that should be a special option to enable dangerous non-deterministic if-you-made-a-mistake-we-will-delete-parts-of-your-program type behavior. I wouldn't consider that optimization though, more like disabling all your compiler's safety features.
People complain about dead code elimination all the time when we have these discussions.
Inlining break code that try to read the return address off the stack frame or that make assumptions about stack layout.
Loop unrolling might change the order of stores and load, which is visible behaviour if any of those traps.
I assure you that for each optimization, no matter how trivial, it will break someone code
If we don't know what will happen that is Undefined Behaviour.
The contradiction you have within yourself is that you know what you want to happen, but that's not what the specification says. If you want specific behaviour you need to specify what it is - not mumble and make a vague wave of the hand about "behavior you know that the target machine will give you" when you've no promise of any such thing. That would come at a cost, and of course you don't want to pay that cost, but that means you can't have what it buys.
Implementation-defined behaviour is a thing. Not knowing what will happen is not an accurate description of undefined behaviour. What the compiler does is assume that undefined behaviour doesn’t happen. When it does happen, it results in a contradiction, and logically every sentence is a consequence of a contradiction (see e.g. “Bertrand Russell is the pope”). That produces all those infamous bugs. Because just like every sentence is a consequence of a contradiction, every program state can be a result of UB. This is untenable.
That is an incorrect assumption, as it clearly does.
It is also incorrect given the standard text.
That’s my entire point. Compiler is free to make incorrect assumptions.
> It is also incorrect given the standard text.
According to C11 standard, section 3.4.3, the standard imposes no requirements on undefined behaviour.
In fact, in the rationale of the original spec., I remember reading that the C standard was expressly designed to be a minimal spec, and that just being compliant with the spec was insufficient for the resulting compiler to be fit for purpose.
And of course the original spec did specify a range of acceptable behaviors, and that language is, in fact, still in the standard. It was just made non-binding. However, it is still there, and pretending it is not seems disingenuous at best.
I agree, but that includes GCC and Clang. ¯\_(ツ)_/¯
And even more critical: Signed integer overflow should be implementation defined, and each implementation do something sane (different from assuming it doesn't happen). This would have saved us many security vulnerabilities, and unnecessary program crashes.
Can't you get that behaviour with -O0 or similar?
Beyond that, you've got a contradiction in your statement -- you can't "force correct behaviour" from a compiler at any point. The compiler always tries to generate correct behaviour according to what the code actually says. If you lie to the compiler, it'll try its best to believe you.
C compilers are intended to accept every correct C program. But they can only do this by also accepting a wide range of incorrect C programs -- if we can prove that the program can't be correct then we can reject it, otherwise we have to trust the programmer. Contrast this with Rust, where the intent is to reject every incorrect Rust program. Again, not every program can be clearly judged correct or incorrect, but in this case we'll err on the side of not trusting the programmer. Of course, "unsafe" in Rust and various warnings that can be enabled in C mean you can tell the Rust compiler to trust the programmer and tell the C compiler to disallow a subset of possibly-correct but unprovable programs, but the general intent still stands.
So if you want to write in a language that's like C but with "correct behaviour" then ultimately you'll have to procure yourself a compiler to do that. Because the authors of the various C compilers try very hard to have correct behaviour, and just because you want to be able to get away with lying to their compilers doesn't magically make them wrong.
The reasoning is something like "if we assume UB doesn't happen, but it does happen, the resulting behavior is unpredictable. This is allowed by the standard, though, because UB allows for any behavior, including that produced by assuming UB doesn't happen."
In other words, major implementations treat UB as preconditions. Violating those preconditions gets you Interesting Results (TM), but that's allowed by the standard because "unpredictable results " really means unpredictable results.
For example, null pointer dereference is UB. If an implementation assumes null pointers can never be dereferenced, it can better optimize some code. If it turns out a null pointer is dereferenced, the argument is that whatever happens then is still permitted by the standard as the standard does not define any program semantics for programs containing UB.
undefined behavior
behavior, upon use of a nonportable or erroneous program construct or of erroneous data,
for which this International Standard imposes no requirements
NOTE Possible undefined behavior ranges from ignoring the situation completely with unpredictable
results, to behaving during translation or program execution in a documented manner characteristic of the
environment (with or without the issuance of a diagnostic message), to terminating a translation or
execution (with the issuance of a diagnostic message).
EXAMPLE An example of undefined behavior is the behavior on integer overflow
https://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdfThe important wording here is "this International Standard imposes no requirements"
i.e. an implementation is allowed to do literally anything in the case of undefined behaviour. It's not quite that the compiler writers are saying "this can never happen", it's more along the lines of "if this does happen, we can do anything at all, including acting as if the conditions were such that UB could not have happened."
So if you multiply two signed ints that the compiler knows are positive, the compiler can assume that the result can't overflow. Because, if it does overflow, the compiler can emit code that does absolutely anything in that case - including acting as if it didn't overflow. Therefore, it can elide checks for a negative result, because either there was no overflow in which case the check is redundant, or there was but the code is allowed to do anything at all - including not performing the check for a negative result.
I'm not sure if I agree with your interpretation of the spec, but even if that's the technically correct interpretation, arguing that things went wrong because the compiler miscompiled the program and that it didn't do the things it was supposed to before it was allowed to do literally anything... just isn't an interesting argument. Things went wrong because your program was wrong.
You might want a language that restricts this, but C is not that language. Or you might want a language that defines all its behaviours, but C is not that language either. Compiler writers put a lot of effort into making their compilers do exactly what the programmer tells them to do.
My personal take is that the correct response to the difficulties of ensuring your C program doesn't exhibit undefined behaviour is probably to avoid writing new code in C. But if you do still need to write C for whatever reason (which I do, occasionally) then it's only sensible to take as much care as the language design expects programmers to take: the compiler trusts the programmer to only attempt operations with defined results.
Honestly, I think we'd all be better served by pushing the concept of "undefined behaviour" a bit further into the background. C has defined behaviours, and the standard helpfully makes explicit which behaviours fall outside the definitions. If you want a defined output then your program had better have a defined behaviour when presented with your input.
I'm not suggesting this is ideal -- far from it, I avoid writing new C code. But it's what C does. If you want to avoid needing to make sure that your program only attempts defined operations, switch to a language that doesn't impose that requirement.
No it's not. Or let me rephrase that. The standard says the following:
Permissible undefined behavior ranges from ignoring the situation completely with unpredictable results, to behaving during translation or program execution in a documented manner characteristic of the environment (with or without the issuance of a diagnostic message), to terminating a translation or execution (with the issuance of a diagnostic message).
Which of these is "the compiler is allowed to assume it does not happen"?
I agree with where you're coming from, but "ignoring the situation completely with unpredictable results" sounds like it pretty much fits the bill. Pretending something doesn't happen sounds a lot like ignoring it completely to me.
What the standard obviously does exclude is nonsense like intentionally reformatting your drives.
Quite the opposite. Assuming it doesn't happen (not "pretending"), is very much not ignoring the situation, at least if you then act on that assumption that it does not happen.
Ignoring it just lets it happen when it does, so if the program specifies an out-of-bounds access, the compiler generates code for an out-of-bounds access, ignoring the fact that it is an out of bounds access.
I'm not sure how assuming UB doesn't happen is distinguishable from a choosing to ignore the situation every time one comes up. You get the same result either way.
For example, a compiler can assume that null pointers are never dereferenced, or every time a null pointer is/may be dereferenced it can just "ignore the situation" with the dereference. I'm not seeing a functional difference here.
This is, of course, subject to the minor problem that "situation" is arguably underspecified. Compiler writers appear to interpret it as something akin to "code path" (so "ignoring the situation" means "ignoring code paths invoking UB"), while UB-goes-too-far proponents appear to interpret it more broadly, more like "the fact that UB can/will happen" (so "ignoring the situation" means "ignore the fact UB will/may happen").
> Ignoring it just lets it happen when it does, so if the program specifies an out-of-bounds access, the compiler generates code for an out-of-bounds access, ignoring the fact that it is an out of bounds access.
Why wouldn't this fall under "behaving during translation or program execution in a documented manner characteristic of the environment" instead?
strcpy(P, filename);
free(P);
if (P[0] == '.') {
// hidden file
// do something
}
Obviously you shouldn't use the above code. However, for illustrative purposes, ignoring the situation [of undefined behavior] probably still results in doing something for hidden files. What people are complaining about is compilers finding that there's UB by static analysis and optimizing out the conditional entirely because they assume dereferencing the pointer to freed memory "doesn't happen."There are good arguments for both sides, in my opinion. But let's not pretend they're the same thing. Deleting logic because it provably would result in UB is not the same as ignoring the UB.
1. This code path is executed. UB will be invoked.
2. This code path is not executed. No UB occurs.
If the compiler "ignores the situation" with UB, code path 1 is ignored (i.e., dropped from consideration). This probably results in the removal of the snippet as dead code.
If the compiler assumes no UB occurs, code path 1 is eliminated as cannot-happen, and the compiler probably deletes the snippet as dead code.
Same result either way (with the obvious caveat that this is one possible interpretation of "ignoring the situation").
> ignoring the situation [of undefined behavior] probably still results in doing something for hidden files.
The problem is that this assumes a very specific definition of "ignoring the situation" which, while understandable, isn't the only interpretation permitted by the Standard.
In addition, there's the fact that such an interpretation would arguably fall under "behaving during translation or program execution in a documented manner characteristic of the environment" instead.
> Deleting logic because it provably would result in UB is not the same as ignoring the UB.
True, but "ignoring the UB" isn't what the Standard says. It says "ignoring the situation", and that's the problem - people can't agree on what "ignoring the situation" is supposed to mean. Compiler writers appear to take it to mean "ignore UB-invoking code paths", UB-goes-too-far proponents take it to mean "ignore the presence of UB".
Ignoring something != killing something. There is no dialect of English in which ignoring something is compatible with eliminating it from existence.
EDIT: Note: I am not personally saying compiler writers are wrong. The body text says it imposes no requirements, so at the very least it's a reasonable interpretation to say the footnote text isn't binding and/or that "possible" behaviors are examples, not a full enumeration. But! On the narrow question of "ignoring" the behavior altogether... Hunting for UB via static analysis and then changing your output based on whether you find it simply is not what the word "ignoring" means.
What you are describing is ignoring the code that has the undefined behaviour due to it having undefined behaviour.
That is not ignoring the undefined behaviour, it is the opposite.
Again, it all comes down to interpreting "the situation". Compiler writers construe it broadly; "Ignore the presence of UB" (i.e., construing "the situation" narrowly) is another possible interpretation, but I don't think it's the one and only definitive one.
In addition, why isn't "ignore the presence of UB" covered by "behaving during translation or program execution in a documented manner characteristic of the environment"? See a null pointer dereference? Just do the "characteristic thing" during translation and emit the dereference. Maybe implementations will need to add documentation somewhere, but that's not exactly the challenging part.
What you are confusing is "being agnostic about something happening or not happening" and "assuming it cannot happen".
And sorry, the "situation" is pretty precisely scoped by "Permissible undefined behavior...". So what can be ignored is this instance of UB, not the fact that UB exists.
Otherwise, if you're going to arbitrarily expand the scope of what the situation is, then how about "the fact that a C spec exist?". That's a situation, after all, and it is the situation you are in.
Or maybe just ignore parts of the spec, like the ones that define what is UB and what is not UB.
Then everything becomes a trigger, and if I can expand scope like that, then I have a standards-compliant C-compiler for you:
int main() { int a = *-1; }
(I am pretty sure you can make a smaller one)I doubt anyone would accept this broadening of the scope.
Once again, ignoring something is not generally the same as assuming it doesn't exist. They could be the same if you assume it doesn't exist and then do nothing differently. However, if you use your assumption that it doesn't exist and act differently based on that assumption than if it did exist, then you are not ignoring the situation.
And the latter is clearly what is happening with today's optimising C compilers. They act very differently in the presence of UB than they would if the UB would not be there, for example not translating code that they would have translated had the UB not been there, had it not been UB or had they actually ignored the UB as the should have.
> just do the "characteristic thing" during translation and emit the [null pointer] dereference
These things overlap slightly, but I doubt that "just emitting the dereference" qualifies as a documented exception to normal processing of a pointer dereference due to UB. It is exactly the same thing it does when the pointer dereference is UB, so it is just ignoring the UB.
Another misinterpretation that seems to be common is to interpret "the environment" in "characteristic of the environment" to include the (optimising) compiler itself.
But it is ignoring the situation? At least, given the broader interpretation of "the situation". I understand there's a narrow interpretation as well.
> And sorry, the "situation" is pretty precisely scoped by "Permissible undefined behavior...".
Maybe? I think I understand the argument. Will need to think on it some more...
> So what can be ignored is this instance of UB, not the fact that UB exists.
I think I haven't been clear enough on this - I had been using "ignore the fact that UB exists" to essentially mean "ignore this instance of UB" - i.e., carry on as if there was no UB. I had been using "ignore code paths with UB" for the broader modern-compiler-style interpretation.
> Otherwise, if you're going to arbitrarily expand the scope of what the situation is, then how about "the fact that a C spec exist?". That's a situation, after all, and it is the situation you are in.
Sure, it's a situation, but I don't think anyone is exactly advocating for an arbitrary expansion of the scope of a situation. "The fact that a C spec exists" is a situation, but it doesn't even pretend to have anything to do with the Standard's permissible UB.
> Once again, ignoring something is not generally the same as assuming it doesn't exist. They could be the same if you assume it doesn't exist and then do nothing differently. However, if you use your assumption that it doesn't exist and act differently based on that assumption than if it did exist, then you are not ignoring the situation.
I'd agree that proceeding without considering UB at all would count as ignoring something.
However, I'd argue that that's not the only way to read "ignore" - dropping something from consideration, to me, certainly sounds like ignoring something. You had to choose to do so, but that doesn't make it not ignoring it. That also depends on framing, though - back to how broadly "the situation" should be read.
> I doubt that "just emitting the dereference" qualifies as a documented exception to normal processing of a pointer dereference due to UB. It is exactly the same thing it does when the pointer dereference is UB, so it is just ignoring the UB.
Sorry, I don't quite understand what you're trying to say with the first sentence - where did the concept of an exception to normal processing come from? The idea was that emitting a dereference is the characteristic translation behavior, so that phrase in the Standard would cover "ignoring the UB" and doing what may otherwise be expected.
> Another misinterpretation that seems to be common is to interpret "the environment" in "characteristic of the environment" to include the (optimising) compiler itself.
I had interpreted "the environment" as including semantics; i.e., the translation environment includes these rules for translation/program semantics, so a characteristic behavior could be "normal" semantics. This interpretation doesn't need to include the compiler since the characteristic behavior is derived from the environment, not the compiler.
Looking more closely at the Standard, though, I'm not too confident in this interpretation. Perhaps "behaving during [] program execution in a documented manner characteristic of the environment" could work, though it's admittedly not what I originally had in mind, and I'm still not sure it works.
----
I do have to admit, though, that after the discussions I've had with you I'm less confident about my understanding of this. It'd be nice to talk to an actual major compiler dev or some C89 committee members about this. Feel like I had run across such a thing at some point, but I don't remember where or when.
> Hunting for UB via static analysis and then changing your output based on whether you find it simply is not what the word "ignoring" means.
Why not? A static analysis pass can flag a code path as containing UB, and future passes can then ignore that path. Sure sounds like "ignoring" to me.
It is fundamentally impossible to both ignore X and make decisions based on X at the same time.
Something cannot be both ignored and a key criterion.
Your argument is akin to saying that a company hiring process that "ignores race" means eliminating applicants based on their race. It is untrue. It is the opposite of true. And I think you know that, so please stop trolling.
Of course you can - the former can be the action you take as a result of the decision. Choosing to not consider something is distinct from refusing to make a decision based on that something, but both are "ignoring" that something.
Again, this boils down to how "ignoring the situation" is interpreted. I can ignore situations with UB and proceed as if those situations aren't present, or I can ignore situations with UB and proceed as if the UB weren't present. The Standard's wording does not rule out one or the other.
In addition, why would the "ignore the existence of UB" not fall under "behaving during translation or program execution in a documented manner characteristic of the environment"? That seems to match "pretend the UB were not present" much more closely.
> Your argument is akin to saying that a company hiring process that "ignores race" means eliminating applicants based on their race.
No, it means that the hiring process makes decisions without considering what race-based effects that may have. If that happens to result in weird race-based outcomes, then that's what happens.
What is doing, in the extreme case, is ignoring it and not generating asm statements for it. Then again, how it could? Code that will trigger UB has no meaning so the compiler wouldn't know what code to generate.
Of course a compiler could assign meaning to some instance to UB.
For example I'm pretty sure that in GCC dereferencing a null pointer is not UB, but it is expected to trap (because POSIX) and the execution not to continue except via abnormal edges (exceptions or longjmp). This means that any code that can be proven to be reachable only through a nullpter dereference is effectively dead code, so in practice it can still introduce bugs if it didn't trap.
The normative text that specifies the semantics of undefined behavior is:
behavior [...] for which this International Standard imposes no requirements.
2. It used to be normative.
That interpretation is the root of the problem. Compilers authors use it to implement outright user hostile behavior in the name of elusive performance.
Why would you spend resources looking for zero days when you can have a few LLVM contributors plant them in every program "as an optimization"?
You can use the same compiler with a language whose design committee isn't deliberately user hostile, like Rust (where UB-like behaviors in safe code are considered soundness bugs).
It should do this in part because very large volumes of code will be compiled that way due to the inability to detect in advance whether a pointer is null or not, and compiling it differently when it is known that a pointer is null makes for inconsistent behavior and vulnerabilities like this one.
This is the way relatively sane simple-minded compilers have worked for a long time, and no compiler should be allowed to pretend that certain constructs do not have reliable or at least consistent system defined behavior that will ensue if they are used.
Similarly if a compiler can tell that a signed addition will overflow in some cases it should be required to do whatever a signed addition overflow does on that architecture, and for the same reason. Issue the appropriate warnings about non-portable code, but do something reliable and reliably useful.
No it does not. Look in the standard.
Seriously though, I get what you mean but what is the point of having a systems-defined NULL dereference? The semantics of the NULL pointer (in C) is such that you are not meant to dereference it. If you dereference it, you have a bug no matter what the manifestation on a specific system is.
And there are probably actual systems where it would be hard to specify the behaviour. For example MMU-less systems, where you can derefence byte 0 but it might not be statically clear what kind of value is stored there?
Maybe C allows a systems-defined behaviour to be implemented as undefined behaviour? But then the distinction is kinda moot.
C standard distinguishes undefined behavior, unspecified behavior and implementation-defined behavior. The first is invalid and inconsistent, the second is valid and inconsistent, the third is valid and consistent (in scope of implementation).
It is rather hard to require invalid but consistent, as the platform itself may produce inconsistent behavior in invalid cases. So what some people want is something like - "be no more inconsistent that would be naive implementation on given platform", but that is both vague and not really useful. And while some of these undefined-based optimizations seems egregious, some are necessary in order to have reasonably performant code (e.g. keeping variable values in registers instead of propagating them from/to memory after each step).
> If you dereference it, you have a bug no matter what the manifestation on a specific system is.
Sure, if you dereference a null pointer you have a bug. But all sufficiently large programs have bugs, and what happens after you trigger the bug is important. The closer the machine code matches your C code, the more likely you are to be able to diagnose the problem from its symptoms.
It's nice being able to strap your program into GDB and have it breakpoint at the point of a NULL dereference because it triggered a segfault, instead of needing to magically infer why results are wrong anywhere in the program because the compiler decided to inject nasal demons due to some subtle optimization pass.
Null pointer dereference is just a special case of invalid pointer dereference. And that does not have consistent results anyways (may segfault or may just return garbage)
The compiler must produce code that does that because it cannot determine whether a pointer is null in advance in most cases. So letting it do something different when it knows that a pointer is null (due to some sort of coding mistake) vs. when it doesn't know that the pointer is null is gratuitously non-deterministic behavior.
The safe (if somewhat slower) thing to do is to emit the same code regardless, so that a null pointer dereference has the same effect even when inadvertently inserted into the program.
The same operation should produce the same result in the same program, as much as possible anyway, and if not the same result then a similar one. It is not a reasonable assumption that any real world program will never dereference a null pointer. It is helpful that the consequence of doing so be as stable and predictable as the underlying architecture provides for.
Your program now behaves differently after optimization. Before, your program crashed with a stack overflow. Now, your program loops forever.
Should the compiler have not done this? Or is it a bug in your program that you passed an invalid input to this function?
People say "I just want my machine to do what my machine does when I dereference a pointer at address 0, why is my compiler making my program do something different?" Why can't I say the same thing in my case?
Of course implementation dependent behavior is a necessity. Undefined behavior too, on the very few cases that nobody complains about. This isn't one of them.
A byte is now always 8 bits. Integers are always 2s complement with wrapping arithmetic. Nobody is going to make a CPU that doesn't use those (at least nobody that wants people to actually use it) because no software would work on it.
For example RISC-V set the cache line size to 64 bytes because that's what everyone else does and going against the grain is too difficult now (and maybe there was no reason to in that case but still...)
The only thing I can think of that benefits from C's "nothing is well defined" approach is CHERI, but I bet a ton of code needs fixes to work with it.
While there is a "group of C users who are unhappy with aggressively optimizing C compilers", they are not unhappy enough to put in the effort to define and implement alternative semantics.
Plan 9 would be a similar attempt to this "Friendly C", or D language: these are attempts by people who were misguided to believe that the original technology was mostly good, and needed only a nudge in the right direction to fix a few problems here and there to make it perfect... but it turned out that the technology didn't succeed on technological merits, and improving the technology only so slightly isn't going to bring any new audience to the clone.
C has way too many problems, far beyond undefined behavior. None of that is important as long as the most popular operating system on the planet is written in C, and C89 at that...
Can you explain why Algol-68 and PL/I didn't generate any network effects?
C may have survived for over 50 years because of network effects. But it didn't get in the position to have network effects by being "awful inside and out*. It got in the position to have network effects by being considerably better than the existing alternatives. (Better for actually writing programs, not better in any theoretical CS kind of way.)
In many ways, I think this parallels the rise of JavaScript, which was far from the best language in general, but happened to be the best language that would run in any browser - and so as popularity of the web grew, so did JS.
Despite small community and not a lot of people working on Fortran compilers (at least as compared to C), Fortran programs usually still beat C on benchmarks.
And Pascal? -- Well, it has so much better grammar... like, it was specifically designed to be unambiguous LL(1) language.
And these are only the two I can name off the top of my head that would straight-up win against C in almost every respect, or at least draw. And these two predate C.
If you look more closely at the history of UNIX and C, you realize that people who created it weren't guided by some great ambition to make a good product... they were just pricks who didn't like to study what others did before them, and thus invented their own square-wheel bicycle. They then also lacked the insight into how defective their bicycle was, but were really eager to sell it to those who knew even less about bicycles. It was through pure luck that UNIX took off and won the OS race. It has nothing to do with its engineering qualities.
Sure... for the kinds of programs that were written in Fortran, which tended to be math-heavy. But nobody wanted to write text-processing programs in Fortran, or parsers, or operating systems, or memory managers. (I mean, seriously, think about writing "grep" in Fortran. You might be able to do it, but C, for all its flaws, is still a far better tool for that kind of task.)
Pascal... better grammar. Horrible to actually use, though, at least before the Turbo Pascal extensions. Text processing was incredibly painful, because there was no such thing as a variable-length string, which was a crippling limitation. I/O was also pretty broken. It was very much not better than C.
Neither Fortran nor Pascal would straight-up win against C for general-purpose programming, still less for system programming.
> they were just pricks who didn't like to study what others did before them
Feel free to cool it with the ad-hominems. They are against site rules.
As someone who has to look into the code of utils-linux, which is very much written in C, I can tell you that C... well, shouldn't have been used there. And the few, but still a significant number of times I had to make a trip into Linux kernel code, I can confidently tell you that C is a bad choice for that kind of program too.
There aren't good programs for C, or, to put it differently, C is not a good choice to solve any problem.
Just to give you an example of a bug in utils-linux I faced very recently, to, hopefully illuminate the problem further: there's a utility called mdadm (short for multiple devices admin). "Multiple devices" is a Linux name for RAID (basically, with minor differences). So, this utility must talk to the kernel a lot, and especially the drivers such as raid0, raid1 etc.
The thing is, and due to historical mishaps... in the previous iteration of this communication protocol tools were expected to use various ioctls to talk to kernel. The system grew and grew, and had outgrown itself. ioctls don't cut it anymore as they need to carry too much information, much more than is plausible to stuff into the simple mechanism that they are. So, the new wave of the kernel-utils communication is through sysfs. But, here comes the horror of every C programmer: parsing! If you talk to sysfs, all you get is file streams. These file streams have structured data in them, but there's no unifying format (so you cannot piggy-back on someone's hard work on creating a universal library to parse that stuff), nor are there any decent utilities in C to deal with extraction of data from file streams.
The result? -- the authors of mdadm discovered that in some circumstances using ioctl isn't going to work anymore, specifically, when dismantling MD devices, but they also realized that parsing sysfs stuff is just too hard for them and... gave up. There's a "TODO" in their code, has been there for many years, that says that they should be using sysfs... and nothing has been done about it.
I could blame the "lazy" mdadm programmers for it, but really, it's a fault of C. Even in some unappealing language like Python, this would've been a no-brainer...
Would Pascal win here? -- Absolutely. A lot of fears that C programmers have when it comes to deal with strings are non-issues in Pascal. But, wait, Pascal has evolved, where C hasn't. Ada is in many ways a spiritual successor to Pascal.
> Feel free to cool it with the ad-hominems.
What you wrote is an ad-hominem attack, regardless of the site rules, I don't care about it. What you quoted, however, is based on memoirs and first-person impressions from people familiar with the subject. It gives fair and appropriate description of the charters in question. While many of them are no longer with us, the remaining ones aren't likely to dispute the claim.
The community was too small. The grows in the programming field was very rapid. I don't want to say "exponential" because I don't have the actual numbers, but you could see it because, well, you would almost never meet anyone with more than some 5 years of experience in almost any programming field for decades, i.e. in 70's, 80's, 90's... I started my career in the 90's, and, so far, in real life, I only met three programmers who started more than 10 years before my time.
The new generation of programmers simply didn't know much of what the previous generation did, and this repeated many times over, not just with Algol or PL/I.
> C may have survived for over 50 years because of network effects. But it didn't get in the position to have network effects by being "awful inside and out".
I don't know how much do you know about UNIX history, but you are willfully misquoting me. Literally, the success of UNIX, and by proxy, of C is the network effect. In a more literal sense than you probably imagine. The appeal of UNIX was more or less this: the "real" computers of the day, the so-called "big iron" had always custom-made operating systems for them. Essentially, every different hardware model would ship with a different OS. UNIX was the first portable OS in a sense that you could install it on more than one CPU architecture (not from the very start, but that was the goal, and they succeeded at it). And the reason why people wanted a portable OS was the idea behind how they wanted to build networks back in those days: have a "real big-iron" have a "side-kick" computer that handles the networking issues. The side-kicks would all run the same system, and serve as adapters to "real" computers. UNIX, in a sense, was a glorified router modem... at least, that's what it was meant to be.
Quite soon programmers realized that instead of connecting "side-kick" computers to "real" computers they may make "real" computers run the same standardized OS. A much simpler one at that! Very little thought was put into thinking about why those systems on "real" computers had to be so complex and big. The same kind of enthusiast who proclaimed that Emacs is huge, but ended up using Eclipse, or the same enthusiasts who proclaimed that Ada is huge, but ended up using C++ are the shortsighted programmers who promoted UNIX in its early days, fullheartedly believing that complexity will somehow evaporate, that they are getting a simple tool to solve complex problems...
So, yeah, because UNIX did became popular due to network effect, literally and figuratively. C just piggy-backed on its success.*
Perhaps more than you. I started a decade before you, so I was there for more of it than you were. (Not at the beginning, I admit.)
> but you are willfully misquoting me.
Not willfully - that takes intent. What, specifically, did I say you said that isn't what you said, or that was out of context? Having read your reply here, I still don't see what I'm misquoting.
You wouldn't believe it, but things often are easier to judge in hindsight, than when being involved with them...
> willfully
You pretend that I said that UNIX got into its position by being awful inside and out, but what I wrote is that it got into it's position despite being awful inside and out. In other words, you pretend to misunderstand me, and then argue with something I didn't say.
In that sense, it was not exclusively built for a byte-addressable machine with registers of a varying size.
also pointer arithmetic on a byte-addressed machine is different from int arithmetic, so you have to know if something is an int or an int pointer if you want to increment it
from the horse's mouth in https://www.bell-labs.com/usr/dmr/www/chist.html, dmr's hopl ii paper, with the advantage of 20 years of hindsight
> The machines on which we first used BCPL and then B were word-addressed, and these languages' single data type, the `cell,' comfortably equated with the hardware machine word. The advent of the PDP-11 exposed several inadequacies of B's semantic model. First, its character-handling mechanisms, inherited with few changes from BCPL, were clumsy: using library procedures to spread packed strings into individual cells and then repack, or to access and replace individual characters, began to feel awkward, even silly, on a byte-oriented machine.
> Second, although the original PDP-11 did not provide for floating-point arithmetic, the manufacturer promised that it would soon be available. Floating-point operations had been added to BCPL in our Multics and GCOS compilers by defining special operators, but the mechanism was possible only because on the relevant machines, a single word was large enough to contain a floating-point number; this was not true on the 16-bit PDP-11.
> Finally, the B and BCPL model implied overhead in dealing with pointers: the language rules, by defining a pointer as an index in an array of words, forced pointers to be represented as word indices. Each pointer reference generated a run-time scale conversion from the pointer to the byte address expected by the hardware.
> For all these reasons, it seemed that a typing scheme was necessary to cope with characters and byte addressing, and to prepare for the coming floating-point hardware. Other issues, particularly type safety and interface checking, did not seem as important then as they became later.
on a byte-addressed machine, byte pointer arithmetic works fine if you treat the byte pointers as integers and don't do any conversions at dereference time; that's what it means to be a byte-addressed machine usually (certainly in the case of the pdp-11)
it's pointers to larger-than-byte things (ints, pointers, and later floats, structs, and arrays) where runtime conversions rear their head; if you try to not distinguish between ints and pointers to ints, then for *(p+1) to refer to the int after *p (instead of one overlapping it, giving a bus error), you need to shift p left by one bit at dereference time, or two bits on a 32-bit machine (if its memory addresses identify 8-bit bytes, as on the 360, pdp-11, vax, and 8086). no such conversion is required for char pointers
hope this clarifies
Well, I had to look into what was this.
And as expected it refers to C uptake against Pascal on the Mac OS.
Except what replaced Object Pascal was C++, not C, even though C minded folks like to think otherwise.
MacApp was ported from Object Pascal into C++, and Metrowerks also added PowerPlant to the party.
MPW ultimately happened, because a few folks pushed for it.
https://en.wikipedia.org/wiki/Macintosh_Programmer%27s_Works...
https://en.wikipedia.org/wiki/MacApp
https://en.wikipedia.org/wiki/PowerPlant
After Object Pascal, the new kid on the block for Apple was C++, even if Macintosh Toolbox exposed API entry points as C like.
Newton OS (Dylan lost to C++), Taligent and Copland were also mainly C++ based.
https://en.wikipedia.org/wiki/Newton_OS
Also, not to be overly Neoplatonic, but is it not the case that C is basically what a smart person would come up with as a way to write portable-but-not-very-abstracted imperative code on a Von Neumann machine?
Sadly, C went forth not only with hard-to-parse syntax, but also with stuff like undefined behavior, null-terminated strings, unchecked pointer arithmetic for array access, etc.
It was designed as a language for a confident kernel hacker working close to hardware, but ended up as a general-purpose programming language for the entire OS, and it's not the best fit for that role.
C and Unix came on to the scene - and they were much much better than what came before.
Users were meant to use the shell commands + awk and get a whole lot done with that. C programs were meant to be small. How small? Well - how much code would you be able to write using ed?
I think Unix intended for programmers to develop languages for users - use lex and yacc to come up with something to hand off to users.
The Unix operating system assumed a corporate org structure that just does not exist. Genius programmers and highly educated users.
https://youtu.be/tc4ROCJYbm0 <--- this is what they were expecting
If we went back in time to tell them how real corporations would be set up, I think C would have had safe defaults that could be disabled when needed and syntax closer to Go than C. In fact if we told them how things would be in 40 years, they would have settled on the erlang VM for stuff outside of systems and graphics programming.
Had AT&T been allowed to profit from their Bell Labs research projects, and history would have been quite different, as proven when they went after Lions commentary book and BSD, shortly after been allowed again to profit from their research.
By that time, C was already popular outside the UNIX world.
Had AT&T made Plan 9 available in equally relaxed terms, history might have gone differently, too.
For example in Spring 1991, with Linux still just some C code Linus Torvalds was thinking of naming Freax if he got it working, JANET, the organisation providing network access to the UK's universities (via X.25 of course) decides to launch JIPS, an experimental IP network.
JIPS was huge. Why was JIPS huge? Because unlike X.25 you could just download BSD source spin up a Unix with TCP/IP and run everything you could think of or write your own software, you don't need anybody's permission - there's some guy at CERN who has written a "Web browser" which sounds pretty interesting for instance. By the time Torvalds writes his Linux 0.0.1 announcement email, JIPS is the dominant use of JANET and X.25 is on its way to deprecation.
It was all bundled together with this free software I got, 100% legit. It seems to work pretty good, and unless you've got funding from somewhere to use something different I think we should use C / IP / Unix.
1. A version of PL/I that doesn’t suck. From what I can tell it was way too broad and the implementations weren’t great.
2. An Algol that is designed in the context of “represents concepts that map cleanly to lower level semantics”. Which was basically just coming full circle from Algol being a way for computer scientists to have something more expressive than Fortran and COBOL, because people kept implementing Algol and realizing that it was missing things. In my own uneducated view, Algol (and many LISPs) are too structured around the concept of completing an evaluation of a program, which made sense in the computing world when they originated, but became out of date as computers started being used for more than just directly computing things.
3. Had good implementations of several “trendy” or cutting edge concepts of the time like preprocessing, recursion, and most importantly structs. Yeah most of this wasn’t technically new. But the prior art like Algol68 was horribly flawed for other reasons.
Because the underlying system allowed concepts like null-terminated strings and unchecked pointer arithmetic, it was fair game for C. I don’t see C as a “better or worse” thing compared to other languages but something that had/has to exist as a bridge between intriguing-but-flawed/limited high level languages and the more functional but unexpressive early languages that saw adoption outside of computer science.
Of course it didn’t need to be the case that the Unix ecosystem’s userspace was mostly C, but C was a huge step up for its time. Like try reading Fortran, COBOL, Basic, and Algol and tell me you’d rather work with that than C. Pascal was later and not that much better, plus computing was a lot more fractured/expensive and the internet was basically not a thing, so it’s not like one could always just start writing pascal on their Nix or vice versa. Even today Rust and C++ are basically the only things that can replace C in many contexts, and Rust is pretty new.
Also, there are a number of historic syntax quirks that a conforming compiler, at least pre-C23, has to understand. For example, function declaration that lacks return value specification (implicit int) or parameters.
I haven't written a C parser, but a number of other parsers. I can confidently say that C syntax, as a result of historical development, ended up in a place where it is much more annoying to parse than a properly designed language with a LL(1) syntax. If your parser can parse the following, I both congratulate you for your persistence and ridicule you for the statement that this is "not at all difficult". C syntax is annoying at the very least.
typedef struct Foo Foo;
void xx(void)
{
typedef int Bar;
Bar x;
}
void foo(int, int, int);
void bar(Foo Foo, int, int Bar);
baz(Bar, y, z)
Foo *Bar;
int y;
double z;
{
return 1;
}
I don't know what exactly is required from a conforming C compiler, but this is successfully compiled by gcc -std=c89.The preprocessor, don't get me started. The D author, who is active on HN, has both stated that parsers are less than 0.1% of the work in a compiler, and also that the C preprocessor is terrible and he required multiple or many attempts, I think spread over multiple years, to get it right.
I don't think it's hard in practice if you use the right approach. More complex from a theory point of view, sure.
I am serious when I say that the Annotated ANSI C Standard book made this easy to understand. Without that book, parsing C types certainly did not make a lot of sense to me either. It can be found here: https://www.amazon.com/Annotated-ANSI-Standard-Programming-L...
PL/I compilers ran on mainframes.
C is a classic case where good enough was so good that many tries at perfect couldn't unseat it.
I never felt like C was a bad tool in the 80s and at least early 90s - in many cases it was better than the alternatives which would result in slow software or would impose difficult constraints (Pascal string limits and array semantics vs. C's strings and buffers). I never thought much about C as a "kernel hacking" language because I was writing software for MS-DOS where systems programming was calling bios routines or intercepting interrupts and doing unholy things. I guess I just saw C as better than Pascal, compiled BASIC, COBOL and slow interpeted languages. When I moved to Unix, C just was the low friction way...
Realistically, C should have been restricted for writing operating systems and device drivers. It is far too low-level for application programming. But since it was often the only portable high level language, software vendors adopted it for purposes that it was never designed to handle. That was partly the reason that C++ and Objective C emerged.
> That was partly the reason that C++ and Objective C emerged. Also Java and Python.
The switch in what customer base to target was what done more damage to Borland, more than anything else.
Delphi could still be a major language on the PC world at least.
For a while they used improved B versions, what Dennis calls New B, Embryonic C and Neonatal C, regarding the language evolution.
It was initially tied to the rise in UNIX; being the C operating system meant that it had to have a C compiler, even if it wasn't always a part of the vendor install. That's why GNU set out to build a compiler before building an operating system.
Similar to how the microcomputer era spread interpreted BASIC as the language of choice.
If you wanted another language, you'd not only have to actively choose it but you'd have to pay money for it.
Rather, it seems that the C standard tried (& succeeded) from the beginning in making it easy to implement a compiler that generated good code, even to the extent of including register allocation hints to the compiler. This also appears to be why the short/int/long types are defined the way they are - not as types with specific ranges, but rather only that long >= int >= short. This allowed a given implementation to map int to whatever was most appropriate to the target CPU, be it 8-, 16- or 32-bit.
In this vein, wanting to be fast on all targets, I don't think C wanted to assume the presence of an MMU or any hardware required to efficiently detect NULL pointers, so instead just left this as an instance of implementation defined behavior.
C isn't alone in this regard - AFAIK many languages only define the behavior of correct programs and leave the behavior of buggy ones as implementation defined.
I used one of these a little, Saber-C, circa 1990, on a Sun SPARCstation. It was glorious, and several steps beyond the Turbo C IDE on MS-DOS that I'd been using at home as a teen.
One of the big wins of Saber-C, before Purify and the later fancy open source memory checkers, was that it could quickly find memory problems. One of my mentors spent some evenings doing a Saber-C-powered memory-bug-search&destroy blitz through an open source Unix X11 game. (I think our own code had less need of help like that, so less Saber-C low-hanging-fruit fun.)
https://archive.org/details/1988-proceedings-summer-san-fran...
* Just because something is undefined in the standard now, it does not mean that it has to be undefined forever. It might be an oversight, or priorities might shift. If you disagree with a particular undefined behavior and about ways how it can manifest you can raise it with the committee and/or compiler vendor.
* The behavior of a program is just as defined by the implementation now than before a standard existed. If you write code that contains operations where the standard does not define behavior, the implementation might (but they don't have to). Consult your implementation for the compilers you care about. They have flags to enable certain extensions that define the behavior of some operations, at the expense of disabling some optimizations.
Assembly language. See e.g. the ARMv8 Architecture Reference Manuals, where code for an abstract machine is listed right there in every single instruction listing, and a great deal of appendix space is devoted to providing a library of helper routines for this machine, and also specifying things like the virtual memory system using the AM.
C programmers who think abstract machine semantics are bad are completely out of touch with reality. I put them in the same bucket as people who think all CPUs are 32-bit x86. They in turn are like the devs who maintain old COBOL systems, except less useful because backward compat and huge emulation advances have made it unnecessary to actually run such 32-bit x86 configurations, unlike the COBOL case.
Anyways, to the actual content of the article, I agree with it but I think the frustration it accepts as reasonable is actually misguided. Here is my understanding of how things ended up for C (disclaimer: I was not born when most of this stuff happened.)
In the beginning, you had a C compiler for your computer, and it was basically just an assembler. This is what the "make C great again" people think the language really is under the hood, by the way. However, very quickly people realized that they want their C code to run elsewhere, and every computer does things differently, so they needed to have some sort of standard of what was approximately the lowest common denominator for most machines and that become the C standard for what it is legal to do. The guarantee created at that point was that if you conform to the standard, every implementation of C has to run your program as the standard specifies. This was palatable to people because they had a bunch of machines with weird byte orderings or whatever, and it was obvious what would happen if the dumb compilers of the time translated their platform-specific code to a new architecture.
Later, the weird architectures started becoming rare. At the same time, though, a new architecture started growing: a virtual architecture, one where the compiler would actually "port" your code to the exactly same processor you were compiling for before, but the code would run faster. It would do this by starting to take latitude through intermediate transformations which it assumed it could do because your program should have been portable to the abstract machine.
Now, this completely weirded people out, because "I'm compiling to a new virtual architecture called 'x86-64 -O3' that is the same as 'x86-64 -O0' but faster and more restrictive" sounds really stupid. It's the same architecture, and they're not even real processors! But if you really think about it compilers are really just taking advantage of the fact that your code is portable, because it works in the space of the C abstract machine, to do a "port" called "run a bunch of optimization passes". People understand when a port ends up causing a trap on another processor because of course it does that on the new platform. But getting people to understand that your unaligned accesses on the "-O0" machine are no longer valid on the "-O3" machine, because, again, the instructions that come out look awfully similar and straightforward most of the time, except for the weird times where a change "surprises" you because the transition between the two crossed through an invalid space. Kind of like a path that seems to have a "weird jump" because it normally crosses through 3D space and at some point someone found a shortcut through the fourth dimension.
Anyways, the performance virtual architecture is all well and good, but what I think will be interesting moving forward is the security virtual architecture, where overflows and out-of-bounds accesses and type confusions are focused on more. Right now as a side effect of performance optimization they end up causing headaches for people, but Valgrind/sanitizers are an interesting look into what compiling to "x86-64 for security testing" architecture looks like. The logical next step is even more exciting, because we're actually starting to deploy real architectures with security-focused features that will require ports that are every bit as concrete as any other physical architecture difference, which I think will "legitimize" this mindset to the people who I called deluded at the start of this now very rambly comment. Page protections means that "const" is not something you can ignore. Pointer signing can mean that your "but they're the same bits underneath!" type confusions are no longer valid. ARM's Morello now means you can't play fast-and-loose with your pointers anymore; they're 128+ bits and you can't just decide you want forge one out of an integer anymore without caring. Ports to these architectures absolutely rely on the existence of a C abstract machine, which has served pretty well considering that its existence is really just what a piece of paper says is legal or not, rather than something really planned beforehand.
https://www.amazon.de/-/en/Robert-Berry/dp/0333368215/ref=sr...
Then there were Small-C and BDS C, as yet two other major subsets in those early days.
[0] - Similar in ideas to Ratfor but applied to a K&R C subset
But in summary, doesn't that just move the target from "The compiler is stupid, it shouldn't be doing this, it clearly should know what I mean and I didn't mean that!" to "This is a bad minimum common denominator, if any architecture really needs this guarantee or for things to behave like this, then they should pay a performance penalty. We shouldn't all have to pay the portability price for this one thing that isn't a issue anywhere."
And to be honest, most of the UB hate I see is about the latter, not the former, no?
The proliferation of optimization passes outpaced decent debuggability, and that's really the problem. Rants against UB are nearly always irrelevant or even just outright wrong. And worse still, those crusaders are harmful. You can see this in Rust as a perfect example. Signed integer overflow is defined two's compliment, much rejoicing from the "UB always bad!" crowd. Except wait a minute, in a debug build of Rust it's defined to be a panic. Why? Because signed integer overflow is 99.999% of the time a bug, and defining how it overflows doesn't actually help anyone. So instead you're left with the worst of both worlds - you both can't rely on how signed ints behave in Rust as a programmer because they have 2 extremely incompatible defined behaviors, and the optimizer/runtime then can't take advantage of them being undefined behavior in practice in release builds to optimize better.
The linters and compiler security flags are there, the problem is getting them adopted.
C libraries like OpenSSL reflect what's culturally appropriate in that language, so even if you came to C from a language with a different culture, too bad it has the culturally appropriate API design and behaviour.
A clear example of this is OpenSSL intentionally mixing uninitialized memory into its randomness pool (because on some obscure and long forgotten platforms it was the only way they had to get any 'randomness'), resulting in any programs written using it absolutely spewing valgrind errors all over the place. (Unless your openssl has been compiled with -DPURIFY to skip that behavior, or had the debian "fix" of bypassing the rng almost completely :P ).
MD_Update(&m,buf,j);
Kurt Roeckx found this line twice in OpenSSL. Valgrind moaned about this code and Kurt proposed removing it. Nobody objected, so in Debian Kurt removed the two lines.
One of these occasions is, as you described, mixing uninitialized (in practice likely zero) bytes into a pool of other data and removing it does indeed silence the Valgrind error and fixes the problem. The other, however is actually how real random numbers get fed into OpenSSL's "entropy pool", by removing it there is no entropy and the result was the "Debian keys" - predictable keys "randomly" generated by affected OpenSSL builds.
I haven't seen OpenSSL people claim that the first, erroneous, call was somehow supposed to make OpenSSL produce random bits on some hypothetical platform where the contents of uninitialised memory doesn't start as zero, it looks more like ordinary C programmer laziness to me.
> I haven't seen OpenSSL people claim that the first, erroneous, call was somehow supposed to make OpenSSL produce random bits on some hypothetical platform where the contents of uninitialised memory doesn't start as zero
I had an openssl dev explain (in person) to to me when I complained about the default behavior: that there had been platforms that depended on that behavior, that they weren't sure that which ones did, and so it didn't seem safe to eliminate it. (I'd complained because I couldn't have users with non -DPURIFY openssl code run valgrind as part of troubleshooting). IIRC the use of uninitialized memory was intentional and remarked on in comments in the code.
- If the "uninitialized" data is actually somehow some kind of interference.
- In LLVM, using a "undef" value will not always do the same thing each time; however, the "freeze" command can be used to avoid that problem. (I don't know if this feature of LLVM can be accessed from C codes, or how the similar things are working in GCC.)
- If the code seems unusual, then you should write comments to explain why it is written in the way that it is. (You can then also know what considerations to make if you want to remove it.)
- Whether or not there is uninitialized data, you will need to make proper entropy too, from other properly entropy data.
One of the difficulties I've had with the 'safer systems programming languages' advocacy is that since something going wrong is inherent and unavoidable -- since the flaw is ultimately in the user's code -- there is a tendency to pretend that the panic isn't something going wrong. In my experience this has result in measurably lower quality code from these communities, code which panics in slightly unexpected conditions -- while something written in C would not (yet may fail in a worse way when it does fail).
I don't think I've yet managed to download and run anything written in rust where it doesn't panic within the first 15 minutes of usage-- except the rust compiler itself and firefox (though I do now frequently get firefox crashes that are rust panics).
It may well be that the increased runtime sensitivity to programmer errors in these languages inherently mean we should expect more runtime failures as previously benign mistakes are exposed, and ought to accept that software written in these languages may be less reliable on aggregate because when it does fail its less likely to create security problems and that this is a worthwhile tradeoff. (Python users sure seem to survive a near constant rate of surprising runtime failures...)
But to the extent and so long as language advocates pretend that panics aren't failures they can't really advocate for the trade-off, advance better static analysis to reduce the gap, and will continue to seem fundamentally dishonest to people who try to use the languages and software written in them and experience the frequent panics first hand.
In principle most panics are amendable to retrying the operation which would be the equivalent of catching exceptions in Java. So yes you get an "error has occurred" warning but your program doesn't terminate immediately. I don't think C has an edge over Java here.
A bit like having warnings as errors, or deciding to ignore warnings as peril to what might come later without the feedback of what those warnigns were all about.
*would prefer if you actually got to pick. But you don't get to pick because once you know of the bug you fix it either way.
Warnings as errors isn't a great example, because if you do it in code distributed to third parties its an absolute disaster as the warnings are not stable and there are constantly shifting false positives. It's perhaps not a good example even without distributing it, because it can lead to hasty "make it compile" 'fixes' that can introduce serious (and inherently warning undetectable) bugs. It's arguably better to have warnings warn until you have the time to look at them and handle them seriously, so long as they don't get missed.
The parallel doesn't carry through to undefined behavior because the undefined behavior isn't logging a warning that you could check out later (e.g. before cutting a release).
I don't understand how this is the worst of both worlds.
You can explicitly define overflow behavior in Rust. There are wrapper types and explicit checked or saturating and wrapping operations if those are necessary for the correctness of your program. If your program doesn't rely on them then checked overflow being the default in debug builds is the way to go and given enough confidence in the final product it makes sense to drop them in release builds and given enough processor advancements we can also do checked overflow in release builds.
In reality that is a trap; C makes you think that you might be able to poke memory in all sorts of ways but in reality there are lot of subtle restrictions around memory access. The whole discussion around pointer provence is the tip of the iceberg here.
And that is pretty much at the core of the whole ub hullabaloo; difference between what C seems to be (or have been) and what the standard says.
If C was just memory, the only operations allowed would be on and through memory addresses, and values wouldn't be first class.
Trivial example of platform where that is relevant is 16b 8086 with anything but tiny memory model (ie. CS=DS=SS). Somewhat more relevant are various Harvard-ish embedded platforms with separate RAM/ROM address spaces. In both cases there are C implementations that are used in production applications.
Another reason for this rule (and probably the original one) are various platforms with segmentation based fine-grained memory protection. That means either things like Burroughs large systems or running C on top of some kind of VM while preserving the underlying memory management. And there are C implementations for that kinds of environments (That most of the existing C code is not directly compatible with that kind of implementation is another thing).
I can relate to people who just like programming in C. :-)
Why am I not surprised about this when talking about C.