C will never stop you from making mistakes
thephd.github.io
thephd.github.io
1. People enabled warnings and -Werror because they want high quality code. 2. Standard can't add warnings because people use -Werror
This means that not adding warnings is directly against the original reason to use -Werror in the first place! We are now avoiding warning people about dangerous things because they requested to be warned about dangerous things!
Which sounds like bullshit to me.
One my favourite warnings in this regard is gcc's misleading-indentation warning. The warning makes sense for new code written by a human, but if the code is machine generated, or decades old without showing any signs of problems caused by a "misleading indentation", then it is indeed much less risky to simply suppress that particular warning in that particular source file or library.
Take your "misleading indentation" warning. If you choose to ignore that warning, you're setting yourself up because I can make a great argument that you don't care that the indentation is misleading. And in fact that you're ignoring the hazards of following misleading indentation which is that another person reading your code could misread it and introduce a defect. And that in fact your policy is to allow some defects, including a possible defect which has killed the plaintiff.
And frankly, the idea that people writing software where defects could kill people would prefer not to be shown new defects because fixing them is an inconvenience is a pretty insulting view of that industry's professionalism and ethics.
And then you might reply "Well no because this particular warning would require us to change some code that's really hard to change correctly so instead of spending the time and expense eliminating a potential defect we just left it in."
Aka. the no brown M&Ms (not the rapper!) policy
https://www.npr.org/sections/therecord/2012/02/14/146880432/...
The fact is that building such medically critical software (which I have done) has many more considerations than warning levels of c compilers.
Additionally, there is the C language and what the standard says are two different things.
Finally, the vast area of standard specified undefined and implementation defined areas substantially impact writing correct software.
Say you've proved memory safety on your source. What you're compiling is no longer that source you have a proof about.
That's what suppressing warnings tend to do, yes.
> but also introduces a potentially semantic-breaking process into your compilation pipe.
Only if the beautifier is broken.
> What you're compiling is no longer that source you have a proof about.
It sure is, again, unless the beautifier is broken.
The point I was making was in context of a discussion focused on mission-critical system. In that context, you can't just add a beautifier to your compilation pipeline with the argument that "the only way things will go wrong is if the beautifier is broken".
Some of the "very influential companies" we're talking about are in aerospace, the automotive industry . . . the LoC numbers are truly immense, the standards compliance rules are incredibly strict, and everything moves slowly. I've never worked in an industry like that.
A more charitable view of the situation would include some representative of an industry like that on the standards committee feeling their heart skip a beat because they realize that what is being suggested would cost millions to implement.
-W2019q2
Cool, huh?The acceptance criteria is controlled by the purchaser. The purchaser can and should audit for this attempt to slide in out of spec code.
struct Meow* p_cat = (struct Meow*)malloc(sizeof(struct Meow));
struct Bark* p_dog = p_cat;
> Most compilers warn, but this is standards-conforming ISO C code that is required to not be rejectedBollocks. That is a constraint violation, ISO C requires a diagnostic for it, and ISO C allows that diagnostic to be an error. The constraint is in the section "Simple assignment", which contains "One of the following shall hold:" followed by a list detailing when assignments are valid. Pointers to different structure types on the LHS vs the RHS are nowhere in that list.
gcc, however, isn't C, just a popular compiler for it, and clang might error on this particular one, but my memory is fuzzy. Someone else might want to chime in here.
Only because it is misquoted, the sentence ends:
> this is standards-conforming ISO C code that is required to not be rejected unless you crank up the -Werror -Wall -Wpedantic etc. etc. etc.
The full quote explains his thought: "Yes, two entirely unrelated pointer types can be set to one another in standards conforming C. Most compilers warn, but this is standards-conforming ISO C code that is required to not be rejected unless you crank up the -Werror -Wall -Wpedantic etc. etc. etc."
Unless you make warnings errors (-Werror), it probably will warn, but will not reject the code (fail to compile).
I already wrote exactly where the standard explicitly forbids it in the message you replied to.
> both operands are pointers to qualified or unqualified versions of compatible types
Not sure, but I still would say he was correct. The compiler might not be able to check if the types are compatible, so the warning as a standard behavior is anticipatory. Sensible, yes.
It's common to get a data payload of struct A, but cast it's pointer to struct B which is a subset of struct A (at is the first elements.
Actually in the Win32 API this is done a lot. Often you're even given a void* to some arbitrary data that you need to figure out what should be there. Sometimes it's not even a pointer it's two integers packed together, in the case of WM_MOUSEMOVE messages.
A* -> void* -> B* is undefined
A* -> B* is undefined
If you want to get the first element, you want, &(A->first), not a cast. The cast isn't guaranteed to be defined, and isn't guaranteed to be the same.
struct Meow { struct Bark dog; }Those are not the same things in C; I think if you go look at your C standard for the constraints on initialization, they are different from and weaker than those for assignment.
If a programming language evolves to the point that previous programs written in that language no longer compile then it's no longer the same language.
So let's keep C as C and, as you point out, new ideas and concepts that would break things can be implemented in new languages.
C does evolve but it's also taken the pragmatic approach not to break the huge existing code base it has. In many cases these programs have been running for decades, they work.
I don't really buy that you cannot ever introduce breaking changes. That's a recipe for disaster and IMHO short-sighted.
>In many cases these programs have been running for decades, they work.
And in many other cases, they have been full of security holes that take an incredible amount of work to discover.
Something like Rust's editions seems to solve this problem very well, and is indeed what they are intended for.
You can still compile old code, you just won't have full access to new features if you choose to do so. This provides a way to keep legacy systems alive while providing an upgrade path for them at the same time.
>I don't really buy that you cannot ever introduce breaking changes. That's a recipe for disaster and IMHO short-sighted.
It's also empirically true in my experience. See Python2/Python3, Perl/Raku, C++98/C++11/C++20. In the latter case, I'm not even sure there are any significant breaking changes, but the feature sets are so different that I find that I pretty much always have to specify which version of C++ I'm talking about.
For such a tool to be effective it must be possible to statically analyse code to a reasonable degree, and C is certainly not a language that enables easy static analysis.
I actually started using C for my side projects since two years precisely because I want very long term backward compatibility (that is, being able to leave a program for years without maintaining it, then make a small edit in it and build it with minimum pain). C is perfect for that, and I agree with the sentiment that backward compatibility is its most important feature.
Rust actually takes a very strong stance on maintaining backwards compatibility, and Go's stance is arguably even stronger in most cases. You're implying incorrectly that these languages are just breaking things left and right for no reason other than "innovation!", which isn't true.
The overwhelming majority of Rust and Go code from years ago will compile without problems today. Any code from post-1.0 that doesn't compile today was (inadvertently) relying on buggy, incorrect behaviors that have since been fixed... and even then, not all incorrect behaviors get fixed because compatibility is considered so important.
C is fine for certain applications, but no one should choose it for side projects based on some notion of backwards compatibility, in my opinion. If some company is building a business application that needs "backwards compatibility" in the sense that it can run on all sorts of arcane microarchitectures and operating systems, then sure... C is still a really painful* choice, but it might be the right choice then, or if there's an existing C code base, then it probably doesn't make business sense to rewrite it any time soon.
* yes, having no protection from footguns, no real standard library, no built-in concept of asynchronous code, and very little of anything useful is definitely painful. If C is the only valid choice for a project, then it's the only valid choice, and that's what you have to do. The number of projects where you simply can't use something other than C is diminishing by the day.
https://timidger.github.io/posts/i-cant-keep-up-with-idiomat...
there are other similar complaints around the net.
While the C and C++ code I wrote over 2 decades ago is now finally "out of date" as of 5 years ago. It took a while. Not 2-3 years.
Which is absolutely attributed to the language being new, not a design fault.
C is already "settled"
Of those who do like Common Lisp, many like it because the standard hasn't changed in nearly thirty years (while simultaneously offering features added to the C++ standard just very recently, e.g. a filesystem path abstraction).
I've been able to compile and run C and C++ code from 20 years ago a truly amazing number of times. It's really surprising at how easy it is to work with well written code even if it's decades old.
Will Rust and Go age that way? Maybe. Too soon to tell.
Not if you have used any third party libraries. Dependency management is a nightmare in these languages.
I spend most of my time working deep in the internals of some things that are 10+ years old running even older versions of some highly (and often badly) modified linux kernels. The well written C/C++ projects definitely stand out.
Or Zig? What other potential C replacement are there?
#define MULTIPLY(a,b) a*b
To ruin your month.
MULTIPLY(2+3,4+5) that expands in
2+34+5 (and not into: (2+3)(4+5)).
To have the latter, you should define:
#define MULTIPLY(a,b) ((a)*(b))
https://stackoverflow.com/questions/14041453/why-are-preproc....
#define DEREFERENCE(b) MULTIPLY(=, b)
int thing DEREFERENCE(ptr);But what’s the point he is making. Is there a link to that lecture?
With all of the novel and esoteric programming languages that exist, I wonder why hasn't there been a "I can't believe it's not c/c++" language that breaks these things, but isn't taken seriously enough to diverge completely from the ISO language standards. (for bonus points, with standardized gcc extensions, and something like embedded asm but for compiled languages (like iso standard c))
Except that C is sometimes a knife with another knife hidden in the grip and it you don't handle it just right, the hidden knife will also cut you. (Thinking of libraries/other people's code)
(I know objects are actually just fancy pointers)
int main (int argc, char* argv[]) {
(void)argc;
(void)argv;
struct Meow* p_cat = (struct Meow*)malloc(sizeof(struct Meow));
struct Bark* p_dog = p_cat;
// :3
return 0;
}
Why declare main that way if you are going to discard the arguments?Why cast the malloc? This isn't C++.
Some people prefer to compile C code through a C++ compiler for various reasons, Microsoft has even been recommending this because their C++ compiler isn't quite as terribly outdated as their C compiler:
https://herbsutter.com/2012/05/03/reader-qa-what-about-vc-an...
(not that I agree with the reasons, but for a library it makes sense to not lock out people who prefer to compile C libraries as C++ when integrated into a C++ project).
struct Bark* p_dog = p_cat;
instead of struct Bark *p_dog = p_cat;
That weird affectation of C++ programmers putting the asterisk on the type and not on the declarator, where it belongs, makes my eyes bleed. I somewhat understand the reasoning, but I think it's a gross violation of the Law of Least Astonishment.That misses the point of the C declaration syntax: you write an expression that when used on its own will recover the basic type. So the asterisk goes with the symbol name, because that's how you dereference a pointer.
Further, it doesn't work if you want to declare more than one pointer like so:
int* a, b; /* wrong */
int *a, *b; /* correct */One way to create a pointer type in C would be to declare it using typedef:
typedef int * int_p;
then one can write: int_p p, q, r;
and declare three pointers to int with perfect clarity, whereas using the C++ style, we'd get: int* p, *q, *r;
which is very confusing, or, int* p;
int* q;
int* r;
which is very verbose. I honestly don't know how C++ programmers typically handle this situation.I have seen some code that strikes a middle ground:
int * p;
which is a little more clear, but doesn't address the multiple declarator situation.So why don't C++ programmers use typedef? I don't know, other than I understand Stroustrup doesn't like it (not without reason).
(Edited for formatting and minor clarity corrections.)
The C++ syntax is the same as C in this regard: a declaration has specifiers, and then one or more declarators.
The exception are function parameters, where you have (at most) one declarator.
> why don't C++ programmers use typedef?
C++ programmers do use typedef. For instance:
typedef std::map<from_this_type, to_this_type> from_to_map;
C++ programmers probably use typedef a bit less than they used to, because of features like auto.When a C++ class/struct is declared, its name is introduced into the scope as a type name. Therefore, this C idiom is not required in C++:
typedef struct foo { int x } foo;
that cuts down some typedefs. If you used a typedef for a C++ class that isn't just a "POD", you have issues, because the typedef name doesn't serve as an alias in all circumstances. typedef class x { x(); } y;
y::y() // cannot write x constructor this way
{
} f(int &a);
to mean "by reference" instead of what it should be, which is "get the address of a, and that will be an int" which is , of course, nonsense. // Inexcusable trompe l'oeil:
int& a, b;
// OK;
int &a, &b;
Here, the mistake may be harder to catch, because the expressions a and b are both of type int, either way. // Intent: b is an alias of a.
// Reality: b is a new variable, holding copy of x.
int& a = x, b = a;
I think what you mean is that the "declaration follows use" principle falls apart for C++ references.That is necessarily true because no operator is required at all to use a C++ reference, whereas the explicit & type construction operator is required in the declarator syntax to denote it.
However, it has little to do with the issue that & is part of the declarator and not of the type specifiers.
Declaration follows use also falls apart for function pointers in C, because while int (* pf)(int) can be used as result = (* pf)(arg), it is usually just used as result = pf(arg).
Declaration follows use also falls apart for the -> notation. A pointer ptr is always being used as ptr->memb, but declared as struct foo *ptr which looks nothing like it.
And of course, arrays can be used via pointer syntax, and pointers via array syntax, also breaking declaration follows use.
Declaration follows use is only a weak principle used to help newbies get over some hurdles in C declaration syntax.
It's not really a big deal because things are also always introduced at the latest possible position. Each is also typically given an initializer. I'd consider it suspicious if I were to see C++ that declared 3 uninitialized pointers back to back like this.
And then you get to the codebases where the authors have chosen to embrace auto and type inference... :)
For example, sizeof(x) takes a type. The type argument for sizeof(int) is different than sizeof(int* ) and the results are different. sizeof(int*[3]) is different as well. These are all different types where pointers change the type. It's not the same type with a pointer modifier, there is no such thing.
The verbose one. A few extra lines rarely matters. I think the number of times this has come up in my code base is very very small, maybe a few dozen extra lines across hundreds of thousands.
Typedeffing pointers, especially for the sole purpose of "being less verbose" when declaring uninitialized pointers, is a red flag too.
Yes. It just looks better IMO. Now, east-const vs west-const?
But we're in 2020 and we've learned to avoid to declaring stuff without initializing it at the same time to avoid the stupid mistake of using something uninitialized.
Interestingly enough, testing with clang shows that uninitialized variables get their warning, but uninitialized raw pointers don't.
int* a, b; /* wrong */
int *a, *b; /* also wrong */
/* correct */
int* a;
int* b; int* a, b; /* wrong */
int *a, *b; /* also wrong */
/* correct */
int* a;
int* b;
/* moar correct */
int* a{nullptr};
int* b{nullptr};A ``typical C programmer'' writes ``int p;'' and explains it ``p is what is the int'' emphasizing syntax, and may point to the C (and C++) declaration grammar to argue for the correctness of the style. Indeed, the binds to the name p in the grammar.
A ``typical C++ programmer'' writes ``int p;'' and explains it ``p is a pointer to an int'' emphasizing type. Indeed the type of p is int. I clearly prefer that emphasis and see it as important for using the more advanced parts of C++ well.
I read the first as 'a struct of type Bark pointer named p_dog'.
How do you read the second example in your head?
int *a, b;
declares a pointer-to-int `a` and an int `b`, rather than two pointers-to-int. The better solution to this problem is "don't do that", but C (and C++) programmers have a fetish for terseness.Indeed, Linus Torvalds has a recent rant about people still adhering to max 80-column width code. It's pointless in the day of massive monitors.
It's not pointless unless you're using your massive monitor like you did your tiny one thirty years ago. I use mine like a bunch of tiny monitors, not one big one.
Also, a take from Stroustrup since I found it. https://www.stroustrup.com/bs_faq2.html#whitespace
void *(*foo)(int *);
foo is a pointer to a function which accepts a pointer-to-int and returns a pointer-to-void.( Taken from https://www.cprogramming.com/tutorial/function-pointers.html )
If the cast is not there, it will tell you that there is a type mismatch.
struct Bark * p_dog = p_cat;The whole "declaration follows usage" is just a bad tradeoff. It makes it easier to parse expressions. That's my understanding of why they did it. It makes it _objectively_ harder to read, because some declarations follow this easy pattern of "name on the right and type on the left", while for some other declarations you have to employ the spiral reading pattern (e.g. for function pointers, arrays, with variable declarations being the easiest one).
You know what's easier that all that? Type on the left, name on the right.
int a;
int* a;
int[] a;
unsigned int(int) f;
Notice that you can probably easily guess what that last declaration represents, without having to consult a wise old man. I think you simply got used to how C does it, so now the actual sane way is weird to you personally, but you have to recognize that it's one additional thing _everyone_ has to learn because it's counter-intuitive. Hundreds of thousands of developers had to learn some weird spiral reading rule because 1 compiler writer found it easier to reuse a yacc rule.That's why Java, Go and D changed this nonsense. Java supports both "int[] a" vs "int a[]". They support both to appease everyone, but they went the extra mile to support "int[] a". D changed it to "int[] a" and called it a day and Go introduced this novel [5]int syntax, which is different, but clearly easy to read ("array of five int").
Again, type on the left, name on the right (or in Go's case, name on the left, type on the right - but at least it's not a mixed bag). Once you see it that way (i.e. you give up on declaration follows usage rule), it's not weird, it's not wrong, it's intuitive, and it makes the language easier to learn and use.
I thought this is a matter of preference until I actually wrote a C compiler for fun and have permanently solidified my opinion on this issue.
The asterisk is part of the type, it's not just some random symbol, it's not a type qualifier like 'const' or 'volatile'. It's a type token that builds a distinct type, e.g. "int * * " is a type that spells "pointer to pointer to int", it's not an 'int' type with some flags attached to it.
Let's take the concrete example of 'struct Bark* p_dog'.
A typical compiler will tokenize 'struct', 'Bark', '* ' and 'p_dog' and will group them as ('struct', 'Bark', '* ') to derive the type "pointer to struct Bark" and ('p_dog') to derive the name of the symbol when it adds an entry into its symbol table, in other words - the compiler itself splits that line so that types are the left, and names on the right.
Because C makes that tradeoff,
int *a;
is idiomatic C and int* a;
isn't.For C programming to be pleasant, you have to understand and agree with the philosophy at least while you're writing C.
Go developers changed this in a particularly elegant way. Here's what Rob Pike has to say: https://blog.golang.org/declaration-syntax.
I too gained a better understanding of the problem after going through a compiler-writing exercise.
I wrote C for a long time before I ever even heard of the spiral rule. I do think that the "declaration follows usage" idea worked a lot better before C declarations became so complex. I'm not claiming that the C declaration syntax is wonderful, it isn't. I just think that the C++ style of pretending that int* is a pointer type is misleading, since that's not really what's happening grammatically. Yes, * is a type token, as you say, but it modifies the identifier, not the type specifier. But since the style now is to only declare one variable per declaration, and to always initialize it, then in practice it isn't actually all that confusing.
I have always wanted to write a compiler and I'm sure I would get a new perspective on these matters if I did. Go has done a lot to clean up C's messes. I haven't used it much, but I'd like to know it better.
We have language standards for changes like this like '-std=c2x' or '-std=c89' with GNU's GCC. I understand and accept the matter of avoiding breakage. Furthermore C is inherently weakly typed, contrary to C++ which is strongly typed. Something you probably should not change, because that are basic language features. But the option to set the language standard does exist for this situation, to allow changes which will affect users. So why it cannot be used here?
That is not a critic. I'm sure that have their rationale for that and know more than me.
PS: Some changes will break the ABI, in that cases we likely see a PREPROCESSOR variable or something like that which is more complicated. The GCC people used it for some changes to std::string if I remember correctly.
more seriously, what I'd like to see in C (as a long-time programmer in C) is less freedom around undefined behavior. I used to feel like the biggest mistakes made in C were around pointer bugs, but you can be careful and get things like that (mostly) right. Undefined code is a lot harder to see and avoid without a very deep understanding of lots of small details.
void count(int x)
{
for (int i = 0; i < x; ++i)
printf("%d ", i);
}
The answer to the latter question is of course "yes" - signed integer overflow is UB, so you invoke UB by passing a negative x.Would you like every loop to be flagged as potential UB? I don't think you'd last a single day programming in that C dialect.
It would be especially useful if you could specify your entry point(s), and let it find all cases where user inputs could cause UB.
I think it would be especially helpful for checking the output of compile-to-c languages.
The basic idea is to replace various common UB scenarios with "defined, but unspecified", in order to move the compiler's interpretation of the program closer to the programmer's.
It fixes most of the C mistakes, while still giving programmers tight control over the system.
- The author is someone quite young (undergrad age) who is serving on a C language committee (that I assume is mostly made up of people who are over 40, probably mostly over 50).
- The author not only is donating his own time to C language committee work, but also clearly knows what he's talking about regarding C.
The article came across to me as thinly-disguised frustration/anger that the committee had no interest in making C "safer".
My take away was that the article very much fitted in with all the articles one sees being positive about Rust not just being of academic / hobbyist interest but being a serious contender for a replacement in many industry contexts.
As making mistakes is nature of human, and C will never stop human from making mistakes, then C will preserve human be natural forever!
-Werror=... for specific warnings might be OK in some cases.
Console homebrew. iOS jailbreaking. Android rooting. Those are only some of the freedom-enabling things this and other "insecurity" allows. It's not all bad --- and IMHO it's necessary have these "small cracks", as it keeps the balance of power from going too far in the direction of the increasingly authoritarian corporations.
I always keep this quote in mind: "Freedom is not worth having if it does not include the freedom to make mistakes."
This is about providing better compiler diagnostics. Such diagnostics can't catch every mistake, as long as we require the aforementioned ability to perform arbitrary operations, but they can catch a lot more mistakes than they're catching now.
I am not tied to some definition of perfection but rather the practical. Developer tools help write safer code. If that code is running dangerous equipment this is even more important.
The expectations and quality needs to be raised. ESPECIALLY in operating systems, device drivers and yes code that has been running "just fine" for years.
Complaining that people don't follow the finer details of the standard while at the same time keeping the standard unavailable to the vast majority of C programmers is a travesty. When compiler writers think that its ok, to put "optimizations" in to compilers that remove vital NULL checks, because the the spec says that something may be UB that make no sense, I think to myself, What would they say If they downloaded the latest version of their favorite text editor, only to find that it would format their system disk when ever the user presses Control? When they then reached out to the maker of the text editor, the developers would answer: "Oh, on page 204 in the documentation that costs money to access it says that pressing Control, is undefined, we we are in our rights to format your system disk". Would they be ok with that and think it was fair? Thats how the C standard body is behaving!
NOBODY learns C form the C standard, and that is your fault, so don't complain about people not following it. The fact that you are also fucking it up doesn't help: (https://news.quelsolaar.com/2020/03/16/how-one-word-broke-c/)
Or maybe you're talking about something I don't understand.
(more or less)
That link seems to be broken.
Also see your sibling comment for other version numbers that may be more useful.
http://www.open-std.org/jtc1/sc22/wg14/www/docs/
n1256.pdf - C99 including technical corrigenda up to TC3
n1570.pdf - C11 final draft
n2176.pdf - C18 final draft (password protected; C18 is basically C11 with corrections)
n2478.pdf - latest C2x draft (now typeset by LaTeX instead of troff)
> In C89, undefined behavior is interpreted as, “The C standard doesn’t have requirements for the behavior, so you must define what the behavior is in your implementation, and there are a few permissible options”.
But isn't the sentence
> Permissible undefined behavior ranges from ignoring the situation completely with unpredictable results
in the C89 spec already giving the compiler developers carte blanche to do whatever they want? Anything can be interpreted as "ignoring the situation with unpredictable results".
In your criticism of C99 (where the word "permissible" was changed to "possible") you write:
> In C99 undefined behavior is interpreted as, “The C standard doesn’t have requirements for the behavior, so you can do what ever you want”. [..]. If for instance you have a large codebase like say the Linux Kernel and there is a single instance of undefined behavior somewhere in there the compiler is free to produce a binary that does what ever it wants. It doesn’t have to document what it does, it doesn’t have tell the user, it doesn’t need to do anything.
This exactly fits the C89 spec. The compiler may just ignore that you are relying on undefined behavior and produce a completely unpredictable binary.
With simple compilers, you could more or less define what happens with undefined behavior. For example, on platform X, if you dereference NULL, you get a segfault. On platform Y, you get a zero value.
Problem is these behaviors are hard to define once you start thinking about the edge cases. You can dereference NULL with an offset, and once you do that, you might skip over any guard pages and overwrite a function pointer with garbage and then jump to some random location in memory, and once you do that, all bets are off. I mean, really, really off.
So that leaves you with three options.
1. Abandon the idea of defining behavior in these cases.
2. Insert bounds checks.
3. Refine the idea of undefined behavior into two separate concepts--"bounded" and "critical" undefined behavior.
C has taken all three options at the same time, in a sense. The main standard takes option #1, the Address Sanitizer gives you #2 (more or less), and annex L gives you #3.
Annex L, option #3 above, is in many senses the sane option--but it requires a fairly large amount of work on the part of the compiler writers, and to my knowledge neither GCC nor Clang have implemented it.
I would also like to add that some of the places where we see undefined behavior were inserted into the standard for real-world cases for unusual platforms. For example, the weird rules for pointer comparisons make sense when you think about segmented memory. Relational comparisons (<, >, <=, >=, but NOT == or !=) on pointers are generally used for traversing arrays, and if you have two pointers into the same array, it would be reasonable to assume that they have the same segment (depending on your memory model)... SO, you omit the segment comparison in this case.
> …not only does compilers act as if its unknowable…
I disagree with this interpretation. The compiler is acting not as if it’s unknowable, but as if it doesn’t happen.
For the benefit of readers here who may not know, this blog post has been discussed previously on HN [0].
Personally, I don't find the argument made in the blog post convincing [0], but I'm hardly an authority on this kind of thing. I've copied part of my comment on the other thread here for reference.
In any case, I think there's an argument to be made that it's not the standards committee that is at fault.
----
I'm not really convinced the author is correct in claiming that a one-word change opened the floodgates to optimizations on undefined behavior. In particular I think:
> Careful reading will reveal that the word “Permissible” has been exchanged to “Possible”. In my opinion this change has lead C to go in a very problematic direction.
is a red herring. In my opinion, the actual problematic phrase is this:
> ignoring the situation completely with unpredictable results
which didn't change between C89 and C99.
It all comes down to what "ignoring the situation" should mean. Compiler vendors appear to interpret this to mean "ignore situations that invoke undefined behavior". Programmers who dislike optimizations based on undefined behavior appear to interpret this to mean "ignore the violation that leads to undefined behavior and treat it like conforming code". Who's right? It's ambiguous.
----
(Edited after noticing that you're the author of the blog post. Sorry!)
However the problem isn't just that compilers ignore UB, they actively make use of it! In the example:
if(p == NULL) write_out_error_message_and_exit(0); *p = 0;
The NULL check is removed by the compiler because of the undefined behavior of writing to NULL. Instead of ignoring that the code might write to NULL, the compiler does the opposite, and assumed that p cant be NULL.
I dont want to sound like I'm showing off of or putting down your comment. Your comment is 100% reasonable. Its an example of how the spec is not reasonable, and it doesn't do what any reasonable person would expect.
void foo(int *bar) {
int baz = *bar;
if (!bar) {
return;
}
/* do stuff with baz */
}
which you can see provably invokes undefined behavior if bar is NULL. In the case you provided where the dereference comes after a function call the compiler cannot optimize out the check unless it can prove that control flow returns, which you claim it does not (it exits). noreturn and such are there to improve these cases and make them explicit to the compiler, but in the absence of information it must be conservative as there are many ways for a function call to never return, one of which is exit(3). Optimizing this case would be incorrect (i.e. a compiler bug) and this happens to be one of the reasons why function calls often serve as a barrier to optimizations in C.If the compiler can't prove that the function returns, it's not allowed to assume that it does.
I challenge you to cook up an example on godbolt.org that shows the behavior you claim. But take note of the difference between
and
I think this might depend on how you interpret the standard and/or compiler's actions.
If you believe "ignoring the situation completely" means "ignore precisely those statements that invoke UB and leave everything else intact", then the transformations done by compilers can look like actively taking advantage of UB beyond what might be implied by the standard.
If you think "ignoring the situation completely" allows "inspect the set of program executions, discard those in which UB is invoked, and optimize based on the remaining possible executions", then the mere act of "ignoring UB" can't really be distinguished from "actively making use of UB"; the actions are one and the same.
If you think languages should facilitate type system design patterns that render large classes of application level mistakes impossible, you are just an architectural astronaut falling prey to premature abstraction and unaware that these languages aren’t making your application code safer or more reliable, only more brittle to the inevitable needs to break its core abstractions to solve expanding use cases.
(Yes, I'm aware that there've been a bunch of studies that didn't find decreased "bug density" in open source repos using statically typed languages, but there've also been studies that found the opposite, and in any event the methodology behind all of these is dubious. Example saga: https://hillelwayne.com/post/this-is-how-science-happens/)
The strong claims come from evangelists of those extreme programming paradigms. You should be asking them for proof that consists of more than anecdata.
It’s backwards to say that essentially what is a historically validated null hypothesis with 50 years of development history on its side is “a strong claim” that requires special evidence, while giving a free pass to all the people using little more than blog posts and slick syntax to claim these extreme design patterns are demonstrably better.
If they are so much better, where are all the companies getting free lunches just by switching to these tools? How is it that the entire industry is so irrational that so few companies are willing to switch?
Superior ways of working catch on very fast, just consider the radical adoption of GitHub and no-sql data systems. Why is strict functional programming not seeing that? What mental gymnastics does it require to take as a premise that strict functional programming is “better” yet adoption rates are super low and successes are not proved with data, only anecdotes?
I really don't see your point here. Are you suggesting that because people have been able to write software in C for 50 years that type systems aren't useful? Because it's not like those programs are bug free. It also doesn't take into account how much effort is need to build and maintain that software.
Also not really sure what you're getting at about strict functional programming. This post was just about type checking.
> Neither will Haskell, Rust, etc.
Rust is really in no way a functional language.
As far as evidence is concerned, you’re not going to solve problems faster, safer, cheaper, more reliably or more extensibly with Rust or Haskell than you are with C or Python, apart from some very niche exceptions.
That just an empirical observation, not an opinion.
Why is there no onus on "established" programming languages (whatever that means)? It's not like production software hasn't been shipped in Rust, Haskell, OCaml, etc. Just because C is older it gets a free pass?
> As far as evidence is concerned
I would argue there's plenty of evidence that languages with better type system are valuable. What we don't have is proof, but that's because nobody has figured out how to do the experiment(s).
http://blog.vmsplice.net/2020/08/why-qemu-should-move-from-c...
Most security bugs in qemu have been C programming bugs, like a NULL pointer dereference.
One is impossible, while the other is already applied.
"I tried punching myself and C let me do it. Now I am ranting how C gave me a nose bleed."
If you are using C for projects that are big and abstract then you are THE problem for picking wrong tools for your task. I killed a mosquito with a hammer but it left a hole in my wall => hammers are terrible tools.