Pointer decay is a fundamental mechanic in C. If you use arrays in C, and you pass said arrays as an argument to a function, then you are dealing with pointer decay.
Pointer decay is a fundamental mechanic in C. If you use arrays in C, and you pass said arrays as an argument to a function, then you are dealing with pointer decay.
It will accept any existing C program unmodified.
With a per-file switch, say, #define __FEATURE_WEAK_ARRAYS, it will start to discriminate T* and T[], make it an error to mix them in assignments, or to pass one instead of the other if both the function definition and the function call are in files with this feature on. It will not complain about functions defined elsewhere.
Then, say, #define __FEATURE_STRICT_ARRAYS the compiler will complain about mixing arrays and pointers as function arguments, no matter where the function is defined. It would require updated stdlib headers, for instance.
Additionally, #define __FEATURE_MULTI_ARRAYS would enable syntax for fixed-size multidimensional arrays, Fortran-style. Now uint8[3][2][10] foo; would allocate 60 elements, and access to foo[1][2][3] would involve one memory dereference, not three.
More support would be needed: sizeof, support for slices, safe array copies and length checks. Nothing extraordinary.
Having this implemented would make a terrific master's degree work.
There are also probably a ton of edge cases you haven’t thought about yet. For one thing, your T[] would be a different beast depending on where it appears—if you declare a variable in a block as T[], it’s an array, but it sounds like your proposal has different semantics for T[] in function parameters—it’s a reference type.
void f(int x[]) {
int y[] = {1, 2, 3};
y = x; // is this allowed?
}
I’m not trying to fight over the specifics of your proposal. I just want to illustrate that the language design is a tapestry, and you’re pulling at one of the threads.There are a few proposals I see like this that circle around. This is not the first array improvement proposal I’ve seen for C. There are also lambdas / closures, which are surprisingly untenable in C when you really dive into it. There’s sum types / discriminated unions in Go, and higher-kinded types in Rust. For each of these features, you can find languages which already have these features, giving you all sorts of templates for how to build it, and yet it’s still such a pain in the ass to add these features to the languages which lack them.
This has been studied for decades; there have been many attempts to build a safer version of C. And eventually the C standard committee will have to decide on an approach or lose out to newer languages at an increasingly accelerated rate.
I'd say that C will lose relevance slowly, more and more, as much as Zig will gain relevance, hopefully to the point of becoming the default choice, and having key parts of Linux kernel ported to it. Not Rust, which mostly replaces C++; even though it can venture on the C territory, its not comfortable nor seamless there. Zig is so seamless, it can even compile your C code along the way. It can do gnarly stuff like handling memory-mapped control registers with relative ease, and with much fewer footguns than C.
C is old, and its age shows. It needs to gradually retire, the way Fortran-77 did.
I have felt a bout of nostalgia for the years past reading this.
T value = foo[1][2][3]; // becomes:
foo -> (T **)
(T **) -> (T *)
(T **) (T *)
... (T *) -> (T)
(T *) (T)
... (T)
(T) <-- This one!
This allows for jagged arrays, yay! So useful.This is in stark contrast to a Fortran-style array, which is allocated as one contiguous piece, all dimensions folded up for linear access with one dereference.
I recommend you read up on what pointer decay actually does; it's more complicated than replacing all arrays with pointers!
In particular, the type of an `arr` expression (after decay) is `int(*)[20][30]`. Decay only ever changes the top-level (outermost) type! And the type of `&arr` is even `int(*)[10][20][30]` -- using `&` or `sizeof` prevents decay from happening. Pointers to arrays are rarely used because using decay is more idiomatic (and because their syntax is unwieldy), but they still exist and would be safer than using decay.
You could if you introduced a new type of "safe array".
e.g.
int[] is a traditional C array which decays to int*
int[@] is a "safe C array" which is syntactic sugar for struct { size_t __count; int* __items; }, and as such can't decay
int[] and int[@] would not be directly interoperable, except by converting both to int* – maybe casting an int[@] to int* would automatically extract the __items member.
(The [@] syntax was chosen at random, if you don't like it, pick something else.)
Most things just take a pointer.
Wouldn't be easy by any means but it could be done at the scale of the Linux kernel if anyone cared enough.
But if you are willing to use something-like-C-but-not-shackled-by-backwards-compatibility, then why stop at arrays and pointers? Just move all the way to D or Zig (or even Rust). These are all languages designed (partially) so that you can port an existing C system bit-by-bit over into them.
Many people who can afford that, are doing that, of course. And that's why you don't really hear much about backwards incompatible developments for C. What would be the point?
Says who? Non-backwards-compatible changes are made to language standards all the time. It's not pain-free, but neither is the status quo.
Besides, who cares what the language is called? Change the name if that's what it takes, but stop conflating pointers and arrays. The cost of that has been literally billions of dollars in losses due to buffer overflows over the decades.
And of course there is C++, the most famous attempt to fix C in a somewhat compatible way by adding more features. It is debatable whether this effort resulted in a better language. C++ has all the bits needed to check array bounds by default but chooses not to do so...
The problem is that the huge existing stock of C code is written in C and not rust, zig, D, etc. The same would be true for your proposed "better C" language and any other incompatible iteration of C.
If you can come up with a way to add these guarantees to C without needing significant rewriting, I can assure you that most C programmers would be very interested.
All of these are very different from C. What I'm proposing is just one small change to the existing C language.
> If you can come up with a way to add these guarantees to C without needing significant rewriting, I can assure you that most C programmers would be very interested.
I can pretty much guarantee that they would not because this is easy: phase in the changes. Start by turning array-pointer conflation into a mandatory warning rather than undefined behavior or whatever it is now. Then wait a few years. Then turn it into an error that you can muffle using -C2024 or whatever.
I actually don't know whether array-pointer conflation is required by the standard or if it's undefined behavior (I'm pretty sure it's one or the other). But if it's the latter then you don't actually have to change the standard to make this happen, all you need to do is write a compiler that does the Right Thing. AFAIK no such compiler exists. But there is just no excuse for this:
% gcc -v
Apple clang version 14.0.0 (clang-1400.0.29.202)
...
% more test.c
int main () {
int x[10];
return *(x+20);
}
% gcc -Wall test.c
%Who is going to update these code once the "do the right thing"-compiler become available?
Oh, and the worst part: some of them may already be bug-free due to 15 years (if not more) of people trying to make money by selling exploits to surveillance vendors or who knows. But there's certainly high-impact bugs left. Now what, refactor the code to use the fancy eliminate-spatial-memory-corruption C variant and introduce a few UAF by the way?
I take exception to this. Of course if I write once-test never, copying from google results and trying to hit jira metrics, then any safety feature in the language will filter out some of the toxic waste code I am producing.
If secure code is designed and engineered, like any other secure technical system would, the language used does not matter so much, but, unsurprisingly, needs to be easy to reason about formally.
I did (because I wrote an exploit for the bug after the Google blogpost, for curiosity), the code looks disgusting. The author (one poor guy) does not have the code in an online VCS and instead dump a source tarball every few months (or years). The upstream vulnerable code was fixed months after news outbreak.
My conclusion is if Apple had a choice it won't end up in iOS at all. Clearly, Apple already paid a lot of maintaining cost in this case (fixing bugs before upstream did), but what you asked for is a whole new level.
(If you're going to say "but they could use it if they relicensed all of iOS as GPL": don't be daft.)
Alas going there ignores all the nice undefined behaviour landmines the standard has buried for you.
So CompCert seems to me to aim to help mission-critical software to move away from C, and possibly into Coq/Isabelle/etc., except for the purposes of compilation to machine code.
I tried to download CompCert so I could try it out, but they only have a source distribution and to build it you need Coq and OCaml and a few other things because of course CompCert is not written in C. No one in their right mind would write mission-critical software in C.
That said, what “mission critical software” are you using that is not running on an OS written in C?
> That said, what “mission critical software” are you using that is not running on an OS written in C?
I'm not sure that's relevant?
If you have a piece of mission critical software, almost all the time you run it on an existing OS like Windows or Linux. You don't _write_ a new OS just for your one piece of software.
Of course, that OS had to be written at some point in the past (and is still being worked on). Presumably that writing was (and is) being done by people not 'in their right mind'. But that shouldn't concern you.
The problem with C is not that you can't write secure-ish software at all; the problem is that this is insanely difficult, and that the trade-offs aren't worth it. Especially for new software.
For software that I get from some third-party, like the OS, I only care about its quality (and price); I don't care about the trade-offs and pains the authors had to endure. If they want to use C in the privacy of their own bedroom, that's up to them.
Of course, Linux in 2024 is written in C, mostly because Linux in 2023 was written in C, then 2022, etc all the way back to the 1990s. There's a lot of path dependence. Back in the 1990s C was a more reasonable choice to write your new OS in. Especially if it was a clone of Unix, C's original home and killer app.
Most safety PLC's boot into a hypervisor that boots an OS (Wind River Linux or something) that runs a program that might be your complied config, or runs a program that runs your configuration (eg code you wrote).
So what languages does it seem likely were used for all those extra layers between your code and the CPU?
And I am talking the sort of controllers that supervise LNG plants, large buildings where they might have more than one elevator in any shaft, prevent overpressuring pipelines and creating environmental disasters and so on.
I would be more comfortable personally if I could write a c program and compile it knowing that the compiled code will run on the bare metal, at least then there are not a couple of closed source proprietary layers of abstraction between me and the processor.
Note : in case you wonder what the difference between a regular PLC and a safety PLC is, a safety PLC has a fuckload more diagnostics. For a safety system PLC, faults aren't the problem, it is dangerous undetected faults. A detected dangerous fault will trip to a shutdown immediately and is an availability issue, not a safety issue.
But, guess what language the firmware that does these diagnostics is written in? I don't know, but I doubt strongly it is one of the 5 PLC languages specified in IEC 61131, so that leaves it likely to be C.
Even those criticial OSes that refuse to move beyond C, most likely are using C compilers written in C++.
Read my previous response to you[1]: you clearly haven't worked on systems that would kill people if things went wrong.
True, but I have worked on a system that would have cost hundreds of millions of dollars if things went wrong. And they did go wrong, though we managed to save the asset. So I do have some relevant experience here.
Yes, if you put enough effort into it and deploy into a non-adversarial environment, you can get the odds of success pretty close to 100%. But then you also get the Therac-25 every now and then.
But mainly you get an endless stream of buffer overflows that lets hackers steal people's bank accounts. That's not life-and-death, but it's a significant societal cost nonetheless.
Thats my understanding too. Code is written in high level systems generating C as output. C becomes rather an implementation detail in a hopefully, more or less completely verified tool chain.
And yet, even though C has been the primary language for safety and life-critical software for decades, with billions of lines of code written to control things where failure results in loss of human life, there has been no significant loss of human life due to the C language.
Throughout the 80s, 90s, 2000s and 2010s C has been the primary language used to control industrial machinery that would kill people on software failure, munitions that would kill people on software failure[1], vehicles that would kill people on software failure, medical devices that would kill people on failure ... and out of these billions of deployments, with billions of lines of code, offhand I can think of only one instance where a different language would have prevented 3 deaths.
I'm not saying that C is safe, but it is clear from the statistics that the danger is very very highly overrated. There is a much greater danger in rewriting battle-tested systems just for the sake of rewriting.
[1] An industry I worked in, btw.
From 80->90->00->10->20s reading and writing C seems less and less magical, including for exploit writers. In 10 years exploits might even be written willy-nilly by an LLM. One of the reasons why writing safe and secure code requires thinking few steps into the future.
I'm not sure what your point is.
It isn't always standard C, if that's what you're trying to say.
It's usually not a hosted implementation, but sometimes it is. It's usually done within industry regulated guidelines, but not always.
The fact is, the "not always" bit matters, because the body of C code controlling actions where human lives matter is so large that there is still a substantial body of standards-compliant C code that doesn't kill people!
The claim being made is contrary to the large body of evidence that we have.
I dunno how relevant that is.
The argument was "Irresponsible to use C for critical systems"
The counterpoint is "Despite being the primary language for critical systems, negligible failures have been attributed to the language."
I'm basically saying this: How do you explain both that severe reaction to using C AND the historically negligible failure rate of the language itself?
You CAN get from point A to point B by riding a horse, but why would you when cars are a faster alternative?
But to the point, many failures have been attributed to the language, most of security bugs stem from the C's lack of memory safety.
_Fortunately_ the reason not a whole lot of deaths can be attributed dirrectly to C is the fact that:
- The safety critical sw is has multiple redundancies baked in, including at the HW level that would safeguard against fatal outcomes.
- Safety critical SW is tested intensley. This proves the "common" cases of usage, but in my experience still fails for long-tail events.
- Memory corruption issues would most of the cases "mearly" lead to resets instead of wrong program output.
- Thinking about SW that is deployed in large numbers, if we admit that memory corruption issues happen in very special cases ( see 2nd pct) then the sudden appearence of a bug could _very_ easily be bundled as a fluke instead of a bug and we would probably not be able to distinguish the failure leading to death as being attributed to C. (since " it works fine on my machine" in 99.99% of the cases)
The one thing corporations wanted, above all, was to increase the supply of programmers who won't break everything. This explains the trends towards safety in programming languages and it also explains why OOP became so popular. It also explains the push over the last 10-15 years or so for everybody to learn to code. Anything that is hard reduces the number of potential programmers which is bad for business's bottom line.
It is only negligible for those that don't have to fix CVE issues.
Which is why we have all those ongoing security laws, companies and goverments have finally started to map money burned due to those CVE fixes.
$ cat tst.c
int main () {
int x[10];
return *(x+20);
}
$ gcc -Wall -O2 tst.c
tst.c: In function ‘main’:
tst.c:3:10: warning: array subscript 20 is outside array bounds of ‘int[10]’ [-Warray-bounds=]
3 | return *(x+20);
| ^~~~~~~
tst.c:2:7: note: at offset 80 into object ‘x’ of size 40
2 | int x[10];
| ^With the changes you have in mind, that new "C+" would be much closer to Zig than C. For a backward compatible bounds-checking proposal see: https://discourse.llvm.org/t/rfc-enforcing-bounds-safety-in-...
This basically just associates a pointer and a length via new (and optional) annotations.
The
void foo(int *__counted_by(N) ptr, size_t N);
with SAL void foo(_In_reads_bytes_(N) int *ptr, size_t N);
https://learn.microsoft.com/en-us/cpp/code-quality/understan...But given how much long time ago XP SP 2 was, and how many people actually use them, unless forced at their job, that is quite telling how much people care.
void foo(int ptr[n], size_t n);
with ptr[n] not copying the full array, just that n is the size.you can try it now with:
#include <stddef.h>
#include <stdio.h>
int ptr[6] = {0,1,2,3,4,5};
#define N sizeof(ptr)/sizeof(int)
void foo(int ptr[n], size_t n) { // error: ‘n’ undeclared here (not in a function)
for (unsigned i=0; i<n; i++)
printf("%d ", ptr[i]);
}
void main(void) {
foo(ptr, N);
}
instead of compile-time: #include <stddef.h>
#include <stdio.h>
int ptr[6] = {0,1,2,3,4,5};
#define N sizeof(ptr)/sizeof(int)
void foo(int ptr[N], size_t n) {
for (unsigned i=0; i<n; i++)
printf("%d ", ptr[i]);
}
void main(void) {
foo(ptr, N);
}
The Linux kernel [restrict .n] syntax is just too weird, almost perl-like, inventing new magic glyphs. And deviating from the normal restrict meaning.What feels bad is you could add standard phat_ptr_t to the C library. But they refuse to do even that.
After 40 years of that I think that was a bad assumption.
I also think that with annotations you can fix code mechanically.
You got
void foo(int *__counted_by(N) ptr, size_t N);
That could be replaced mechanically by void foo(sized_buf_t buffer);
And if it can't that's already a big problem. typedef struct {
int *__counted_by(count) buf;
size_t count;
} sized_buf_t;Zig is hands-down a better language than C, and (I'll take your word that) it fills the same niche as C, but it is still a different language with its own idioms and lore and conventions. It is not C-with-tweaks. It cannot be compiled by an extant C compiler. Code written under my proposal would be legal C code under the current standard (but not the other way around).
[EDIT] Actually, that turns out not to be true. You'd need to change the behavior of SIZEOF or provide some other way of getting the size of dynamically allocated arrays at run time, since this information would now be maintained by the compiler.
You can do this:
#include <stdio.h>
int main(int argc, char *argv[]) {
unsigned int lens[argc];
for (int i = 0; i < argc; ++i)
lens[i] = (int) strlen(argv[i]);
printf("Computed %d lens into %zu bytes of array\n", argc, sizeof lens);
return 0;
}
Very contrived pointless example, but still.std::array::at does bound check.
Opinions differ on whether this is the great strength or fatal flaw in C++.
Which most sane compilers will do for you in debug builds.
Sorry, I challenge your authority to decide what is "most natural way".
I feel like the amount of effort that has been spent so far to make C safe and fix bugs due to C not being safe is greater than the effort that would have been required to rewrite all existing C code into memory safe languages.
But I think secretly C programmers don't want memory safety. Dealing with pointers and remembering to malloc and free are part of what makes them feel more skilled/elite than those other programmers who have garbage collection and bounds checking.
Memory management, array bounds checking and a bunch of other 'safe' features have a price that I'm not willing to apply broadly and redudantly to all of my software.
I'm going for speed, that's why I'm using a Ferrari. Corolla's are fast and safe - use those, don't lobby for Ferrari to add safety to their cars at the expense of speed.
There are hundreds of languages. Use those. Write transpilers for C code for software that shouldn't have been written in C because it had to be safe. That would be a better use of your time.
If you’re writing large systems Java is the fastest language. Go benchmark Jetty vs Apache when serving non trivial web apps. Java is actually amazingly fast but it feels slow due purely because it starts up slower, startup time is not an issue for long running applications.
Heck just look at Apache Lucene the gold standard of full text search.
Your comment is confusing "fast enough" with "fastest". Java is fast enough for lots of applications and that's fine, particularly because large systems are usually I/O bound, but it makes no sense to conclude that it is therefore faster than C.
I've been writing Java code for the past 10 years. Do you know how people speed up Java applications? They write the code in C, compile it as a library and use JNI to invoke it.
I would recommend you brush up on your CS fundamentals.
What C does is assume that you're willing to sacrifice correctness to make the compiler simpler which is quite different from what you described.
In practice this has a negative consequence for performance as well as safety.
Array bounds are not being checked on every array access not because it would make compilers too complex.
Correctness is being sacrificed mostly for speed or portability on future CPUs.
There are examples of language features that simplify compiler writing, however.
For example, type promotion from char to int is a feature that reduces the number of cases one would have to deal with when implementing the type system in a compiler, but it's there because it doesn't sacrifice neither performance nor portability.
Other replies in this thread mentioned they have similar problems writing Go, I don't know to what extent it applies, in my limited experience working on Go codebases I never see such issues.
When it does in fact cause an issue with project delivery acceptance testing, the issue is solved by making use of profiling tools, and cirugically disable bounds checking, which most systems languages since the dawn of time also support.
Unfortunately for what I do I had to do this a lot. I guess that's also why I'm not seeing it in Go, never tried to write a query engine in Go.
The algorithms, networking protocols and thread scheduling are much more relevant, that the bounds checking done in the C++ data structures.
As for writing query engines with bounds checking languages, there are several examples.
That might be true, but you could still specify something slightly less exploit heavy than 'undefined behaviour'. Eg you could make out-of-bounds access into implementation defined behaviour.
Undefined behaviour in C infects the whole execution, not just what comes afterwards.
https://devblogs.microsoft.com/oldnewthing/20140627-00/?p=63...
I'm not deep into compiler construction, but to me these examples seem just like a logical consequence of what UB is -- it's a (runtime) situation that the compiler is not required to take into consideration. It can opt to not emit code to treat these situations at all, etc -- effectively assuming they don't happen. The point is to allow the compiler to blindly dereference a pointer even when it can't prove that the pointer is valid. Or to allow it to implement arithmetic on a register of bigger size (assuming the computation doesn't overflow), etc.
Now, depending on how optimizers are written, the compiler may end up inadvertently detecting UB and optimizing out entire branches of code, just by virtue of how the optimizer works internally. You can bet that the compiler doesn't think much of e.g. what is earlier or later in time, when doing optimizing transformations.
Of course a "miscompilation" (of code that is buggy in the first place) is an unfortunate situation and a diagnostic would be better. Compilers should improve (and they probably do). Compilers should be friendly and give unsurprising results and good diagnostics as much as possible.
But to "define it to not travel backwards in time" right in the spec would probably be very hard and might negate the point of UB in the first place. It would require doing the work of compiler authors, which are the people responsible to figure out how to make _their_ compiler solid and ergonomic while also offering the optimizations people want. This is already a hard task for the authors of a specific compiler, and probably not something that you can easily define in a language spec!
And for balance, I've never consciously had to deal with a miscompilation like this, and I write C and C++ in professional capacity almost every day. Instead, most bugs I deal with are of the most trivial kind, you hit a segmentation fault, quickly navigate to the piece of code where there is still some initialization stuff missing, and fill it in. Or there is a logic bug that is entirely unrelated to UB, those are in fact, typically, more difficult to find and fix.
Note that while I'm by no means an exceptional programmer (not that I think you think that of me). I simply want to solve a problem. And while developping I introduce bugs and even UB sometimes (even though it seems to be quite rare if I can trust sanitizers). I'm actually sophisticated enough to develop in debug mode, with most optimizations turned off, and this might be one explanation why I've never hit an annoying situation like this.
To me, these stories are fascinating, and I think they should be taken seriously. But their effect on online forums is mostly to heat up discussions.
But of course the proposed C++ 26 Contracts rely on C++ expressions. In C++ the expressions are themselves full of potential UB footguns, including signed overflow and illegal pointer de-reference. Thanks to time travel, this means adding the "safety" pre-condition may actually make your software much more dangerous not safer.
One proposed way to defuse this somewhat is to prohibit that time travel. Your contract expressions might still be UB but the idea is to promise by fiat that if so this doesn't actually time travel and destroy previously correct parts of the software.
I genuinely don't know what will happen there and can offer no predictions. In terms of what would be amusing as a spectator I hope either SG23 (Safety) explicitly says this is a terrible idea but WG21 ships it anyway or, equally funny, SG23 endorses the current unsafe nonsense as safe and then a subsequent committee has to establish a "Safety but really this time" Study Group to replace SG23 in a few years when it's thoroughly discredited.
> most bugs I deal with are of the most trivial kind, you hit a segmentation fault, quickly navigate to the piece of code where there is still some initialization stuff missing, and fill it in
Sure. C++ is such a bad language that most of your bug fixing is stuff which wouldn't even happen in a better language. Rust's std::mem::uninitialized<T>() is ludicrously dangerous, so it's deprecated (as well as unsafe) and yet C++ not only does this, it's silently the default for the built-in types. Hilarious. My sense is that a correct fix for this won't land for C++ 26 although maybe Barry can get the stars to align and prove me wrong.
Quite honestly I don't recall signed overflow to happen, like ever. It's probably happened at some point but I really don't recall. I'm not trying to make it happen because I don't have a use for it. It's not useful anyway to have a number wrap around from e.g. 2^31-1 to -2^31. It is useful however to wrap from UINT_MAX to 0 (modular arithmetic), and this is in fact defined.
Of course, if you write "if (x < x + 20)" and turn the optimizer to -O7, then the compiler will run the body unconditionally, even though assuming signed overflow the test should fail when x equals INT_MAX. Woah, I'm crushed. That condition is exactly what I needed to write.
> Sure. C++ is such a bad language that most of your bug fixing is stuff which wouldn't even happen in a better language. Rust's std::mem::uninitialized<T>() is ludicrously dangerous, so it's deprecated (as well as unsafe) and yet C++ not only does this, it's silently the default for the built-in types. Hilarious. My sense is that a correct fix for this won't land for C++ 26 although maybe Barry can get the stars to align and prove me wrong.
I mean I could just write "#error Unimplemented" to get a compile time error but I'm not bothering. It seems what you describe as a terrible memory safety bug is simply my way of browsing to the next piece of code that I need to work on. Go figure...
Are you still developing C/C++ code? I get the impression you've given up on it and have jumped on the Rust train a hundred percent. At least there is a huge disconnect between the pictures you paint and my own development experience from daily practice.
But to make it clear again, I'm obviously not opposed to having the compiler issue an error whenever it's able to detect UB statically. In fact, this is how it should be.
You seem very confident how the compiler will react to UB, I wouldn't be. You also seem unduly confident that you can spot such a footgun and wouldn't pull the trigger.
> It seems what you describe as a terrible memory safety bug is simply my way of browsing to the next piece of code that I need to work on. Go figure...
It's Undefined Behaviour, and you're just quietly confident that it'll be fine. Which it will until it isn't one day (and maybe that day was yesterday).
> I mean I could just write "#error Unimplemented" to get a compile time error but I'm not bothering.
A compile time error seems like a weird choice. Why write such an error only to immediately have to fix it? In Rust I'd write todo!() when I need to come back and actually provide a value or write some more code here later, that way it only blows up if this code actually executes.
> Are you still developing C/C++ code?
Not in anger for several years. I write Godbolt-sized samples to make a point sometimes.
> But to make it clear again, I'm obviously not opposed to having the compiler issue an error whenever it's able to detect UB statically. In fact, this is how it should be.
All the popular C and C++ compilers provide a great many flags you can set to get more of these diagnostics you're "obviously not opposed to". How many are you using today? How many did you try and then turn back off because of all the "false positive" diagnostics about things you knew were a bad idea but have preferred not to think about because hey, it seems like it works, right ?
Well that's exactly what I get too by doing nothing and noticing the segfault when running my debug build. Sure, I get it, it's UB and there could be "time travel" and what not. But in practice I seem to get my segfault, so that's just how I end up developing. If it wouldn't work, I could write my own todo() macro, nothing magical about it right?
> All the popular C and C++ compilers provide a great many flags you can set to get more of these diagnostics you're "obviously not opposed to". How many are you using today? How many did you try and then turn back off because of all the "false positive" diagnostics about things you knew were a bad idea but have preferred not to think about because hey, it seems like it works, right ?
I compile with -Wall on Linux and -W4 on MSVC. If I'm not seeing bugs in the integration tests, there is for most domains very little economic incentive to setup various static analyzers etc, so I rarely do that. I run -fsanitize on some of my stuff from time to time just for kicks, but haven't gotten enough value out of it, which is why it's not a habit for me.
But since you mentioned it I went ahead and ran -fsanitize=undefined -fsanitize=address on a test program of the multi-threaded queue I'm working on, which is a bit performance-oriented -- on my older desktop computer it persists > 2M individual messages/sec (600MB/s) to a single disk, with to-memory message submission latencies of < 300nsecs for 99th percentile, < 2usecs for 99.9th percentile and < 30usecs for 99.99th percentile. The test program runs for ~6 seconds, submitting 16M messages (4GB of data), with 4 concurrent readers receiving the messages as soon as they come in. 178 fsync() calls were done by the enqueuer threads or the dedicated flusher threads. There are various internal buffers (a couple MB) and multiple internal message stores (1 optimized for fast submission / 1 for dense storage), and a couple low-contention mutexes but also some wait-free stuff.
-fsanitize didn't find a single UB (I double checked that the detection does work in principle by introducing a signed-overflow bug and a null-pointer dereference as well as an OOB memory dereference). And it found 3 leaks of 1 byte, which seem to be false positives: all related to smaller structures (more than 1 byte) that I allocated and freed correctly. That's all it reported.
I then went on to test using valgrind, which notably reported 0 leaked bytes, and otherwise only reported tons of spam exclusively related to printf-family calls. IIRC these are common false positives due to library mismatches or something like that. You can get rid of them, but I won't bother now.
This is the first time I tried static and runtime analyzers on this project, other than -Wall. In other words, it seems that just by fixing bugs and adding code until it worked, I produced a software of ~5K lines of C code that performs quite well and has 0 bugs or UB uncovered in the good hour of work that I put in.
https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/p32...
In my sibling post I merely want to illustrate that all these concerns have little bearing on my day-to-day work (which mostly doesn't need to be certified, and is not related to the defense industry or similar). Some of these I perceive as FUD, as said I know that you can provoke these situations but I've never personally encountered nasal daemons in practice, and I feel quite productive, am not spending a lot of time on bugs, so why bother.
So, one way to understand these papers is that Microsoft (at least some parts of it) thinks unsafe Contracts are worse than no Contracts. Now, would that mean they just won't implement an unsafe Contracts feature shipped in a C++ 26 document? Maybe. Would these fixes get it over the line? Maybe.
Without being involved -- I have no intention using any of these Contracts in whatever form. I will say though that I wouldn't care if there is UB in the contract language (just like there is in the normal language). I would prefer the variant with UB if it is simpler and more aligned with the language core. Removing the UB here is an academic exercise. Safety absolutists are uncompromising about the goal of correctness and provability. They are blind to the pragmatic issues created by the idealism. Contracts in either form could probably improve correctness by a lot, like 99% or whatever. So why should I care about the paper which could in theory bring the remaining 1%? It doesn't affect me pragmatically.
The flaw with either is that this is only in theory. In practice, I will never create enough formal contracts in to significantly improve correctness. Whatever system it is. Why? The costs are just too damn high, the only way to achieve 100% correctness when considering also pragmatic concerns, is to just not write any code.
My approach of just coming up with a simple design (not in code), trying to implement it in the most straightforward way, and fixing the code until it works, as described in my other comment, seems to have achieved something very close to correctness (maybe even 100%? Probably not).
Again, I'm not saying that UB is good or should be tolerated. I don't want it in my programs and if I find an instance of UB I'll try hard to get rid of it. However there is a reason why UB exists in C/C++ (as well as many other languages that may not have as much of it, but still have a lot of it even when not defined explicitly). And alternative approaches, trying to prevent UB mechanically, come with a cost that may not be worth it depending on what you're working on. I feel strongly like it isn't worth it for me. If you're building a fully verified or certified product, tradeoffs are likely different.
If we're citing big names, here is a well known person describing their view, which I find myself agreeing a lot with.
Do you have suggestions on how to fix the incentive?
If the engineers actually admitted that C is not a safe languages for shipping software, then we could at the very least freeze the existing code and write everything _new_ in a same language. But we don't. Engineers still go starting brand new greenfield projects in C which is just insane.
Sure, if you are willing to help, here's my wishlist:
- I wish we can freeze libssh and write everything new in Rust.
- I wish we can freeze CPython and write everything new in Rust.
- ...
Can you do it for free? At work I'm busy maintaining old projects in C++ and writing new ones in Rust. Since I'm not getting paid to rewrite or maintain our dependencies full-time I can't do above. Oh, I'm not paid to initiate an effort to freeze our old projects either.
If this sounds too harsh:
- I wish we can freeze ZMK [1] and write everything new in Rust (or Zig, though it's not memory safe, whatever).
That's about one of my hobbies and I always wanted to do it.
Having just come out of embedded firmware land: it's not secret, a few members of my team were pretty open about either not caring about or not wanting memory safety. But the added productivity that Nim gave us outweighed their complaints in the end
How much, do you think, would rewriting all existing C code cost?