C23 Implications for C Libraries
gustedt.gitlabpages.inria.fr
gustedt.gitlabpages.inria.fr
I wonder why they’re disallowing _BitInt(1). It would be a signed integer type with the possible values -1 and 0.
Prior to C23 there was no guarantee of two's complement, so it was possible that such a field could only hold the value 0.
It does say bitfields of size zero are allowed (that forces padding to the size of the allocation unit), and also, curiously, says that, in C, “whether int bit-fields that are not explicitly signed or unsigned are signed or unsigned is implementation-defined.”
even better, adding constructor and destructor for RAII in c, main() has it via attributes __constructor__ __destructor__, gcc has __cleanup__ for non-main() functions too, let's add it to C directly?
Constructors are a nuisance, and destructors while great require significant additional semantics to avoid double destruction.
C++ uses unconditional destructor thanks to ubiquitous copying and non-destructive moves, but that requires all objects to always be in a destructible state.
Rust instead relies on lexical drop flags inserted by the compiler if objects are conditionally moved.
They’re more convenient and composable than defer/cleanup/scope, but they also have more langage impact.
Obviously the trade off is defer pushes that issue onto the developer.
I think the Rust view of C++ will turn out to be an over-reaction. The issue wasn't a failure to make the implicit explicitly provable, it was simply to make it explicit at all.
It’s not. It’s a conveniently simple language addition which foists all the issues and edge cases onto the language user. Defer is a limited subset which is more verbose and more error-prone.
However like “context managers” (e.g. try blocks, using, with, unwind-protect, bracket) it also works acceptably in languages where ownership concepts are non-existent or non-encoded, whereas destructors / drops require strict object lifetimes.
> The issue wasn't a failure to make the implicit explicitly provable, it was simply to make it explicit at all.
What, pray tell, would be the use of making borrows “explicit” then ignoring them entirely? Useless busy work?
My point is that requiring compile time proofs is likely too far to reach the practical goal.
Indeed, it seems most Rust code today agrees: widely turning off the borrow checking with copies and the like.
I here this and that `unsafe` is widely used in Rust code bases repeated a lot, but is there any actual data backing this?
The existence of these three "dialects" goes to my point that the borrow checker isnt The Ideal solution to the problem of high-perf. safe static computing.
That link goes directly to the relevant section in the manual.
It includes a cc wrapper called `cedrocc` so you don’t need intermediate files, and it works hard to produce clean code for the generated parts. The rest is not modified at all. The goal is to be useful even if you only use it once to generate that repetitive code.
The pre-processor `cedro` depends only on the C standard library, and `cedrocc` depends on POSIX.
[1]: https://en.wikipedia.org/wiki/The_C_Programming_Language
Secondly, Bjarne's books include the overview of the standard library, which neither K&R nor its ANSI C revision do.
Third, since everyone keeps using compiler extensions in C, moreso than in C++, those should be included as well.
You can have basic safety and QOL features without making the language a monstrous beast. Hell, you can cut old garbage like K&R declarations or digraphs to make room. You can even remove iso646 and most of string.h as a gimme.
`defer` would to some degree solve that issue.
Similarly, I've been missing nullptr, just for the expressiveness. I like that C23 now includes it :)
How can we support/encourage you? (in case you need it)
The argument against "Why dont they just fix X?" usualy commes down to: it would break millions of lines of code, make every C tutorial/book obsolete and force millions of C programmers to learn new things, not to mention that if we broke the ABI, we would break almost every other language since they depend on Cs stable ABI. Breaking C would literaly cost tens if not hundreds of billions of dollars for the industry.
Look at the move between Python 2 and 3. The cost of breaking bakwards compatibility have probably been astonomical, and then there is way less Pyton code being maintained then C code.
C operates on a scale that is almost unfathomable. A 1% perfomance degredation, in C will have ameasurable impact on the worlds energy use and co2 emissions, so the little things really matter.
C is "simple" in the same way DNA is "simple". There is a small set of very straightforward rules, but that set provides immense flexibility. But it in no way means that any resulting object will not be complicated. Perhaps a better word than "simple" would be "non-complex", making the distinction between "complex" vs "complicated" systems.
By contrast, the features you refer to would increase the "complexness" of the language (I won't even dare use the grammatically correct word "complexity" here, lest we deviate into yet another trap of conflating definitions).
Which, I would agree with OP is, for better or worse, probably far outside C's goals as a language.
You can get into pointer and memory errors in C++, Rust and Ada. All of them low-level system languages. Sure, those errors might be harder to produce, but not impossible, and definitely easy enough to still trip you up.
I programmed in all of those languages, except Rust (just don't like it). At least in C you pretty much now WHY (not necessarily where in the code) things went south, without consulting a thousand page specification or having to remember the myriad of language feature interactions that could have triggered those problems.
Moreover, C being small, it's a good on/off language. Try doing a code review for a C++/Rust/Ada code base which uses features heavily after not having touched the language for a year. I bet it is not as easy as C.
You know, some things in life are just hard. And low-level programming is one of those things. C is only honest about this.
Umm, why is typing something like defer (and documenting what it does) different than typing something like goto (and documenting what it does)?
Frankly, not being comfortable with adding safe and efficient abstractions to a language will result in the death of the language (both for spoken languages and also for programming languages).
Because goto is explicitly declaring a control flow where defer causes implicit control flow in code that does not explicity declare it. That's a meaningful difference that the OP was trying to specifically avoid regardless of whether you care about it.
Why is this different than any loop in C? or function without a return? Wouldn’t it be more explicit and very minimal overhead to instead always use a goto at the end of a looping block?
Even tho C allows while loops (as a mistake - all control flow should be explicitly defined at the end of the block), why aren’t you - as a best practice - using an explicit goto and if block instead of while / for?
Because it's non-local control flow. Did you just ignore my whole comment where I pointed out where cases like exit() unwinds the stack, invoking deferred block in each caller up the call chain back to the root? Does for/while do that?
> invoking deferred block in each caller up the call chain back to the root
Yes if the caller writes N defers in the same function or across M functions then the code will run N defers.
Why is this any different than nesting N loops, in the same function or across M functions? It’ll still run N end of block control flow statements without them being “explicit”.
At the end of the day “explicit” typically is really just a replacement for saying “familiar with the previous documentation, don’t change things”.
Anyways, that’s all I can put into this conversation. Please just try to consider that you’re conflating familiarity with simplicity and leverage System 2 to really ask yourself if what you’re favoring is valuable to the group or just the NIMBY.
Defer is useful for code clarity in many situations IMHO, and it's easy enough for a project or code QA tool to ban its use for those that don't like it.
It may be deceptive. Just look at how various compilers or even the same compiler with different options interpret the same C code.
> as explicit as possible
C is not assembler.
> they are much too inflexible because they bind 'behaviour' to types
1. that is the baseline you want, data is not an amorphous and meaningless blob, the default for a file or connection is that you close it, the default for allocated memory is to release it, etc…
2. that also allows much more easily composing such types and semantics, without having to do so by hand
3. and is much more resilient to future evolution, if an item goes from not having a destructor to having one… you probably don’t care, but if it goes from not having a cleanup function (or having a no-op one which you could get away with forgetting to call) to having a non-trivial cleanup function you now have a bug lurking
4. destructors also make conditional cleanups… just work, the repetitive and fiddly mess is handled by the computer, which is kind of the point of computers, rather than having to remember every time that you need to clean the resource on the error path but not on the non-error path, defers easily trigger double-free situations; this also, again, make code more resilient to future evolutions (and associated mistakes)
5. furthermore destructors trivially allow emulating defers, just create a type which does nothing and calls a hook on drop
> Data should be 'stupid' and not come with its own hardwired behaviour
That’s certainly a great take if you want the job security of having to fix security issues forever.
Something like a file or connection is already not just 'plain old data' though.
In my mind, 'data' are the pixels in a texture (while the texture itself is an 'object'), or the vertices and indices in a 3D mesh (while the 'mesh' is an object), or the data that's read from or written to files (but not the 'file object' itself).
C is all about data manipulation (the 'pixels', 'vertices' and 'indices'), less about managing 'objects'.
For objects, constructors and destructors may be useful, but there are plenty cases in C++ where they are not (textures and meshes in a 3D API are actually a good example for where destructors are quite useless, because when the 'owner' of the CPU side texture object is done with the object doesn't mean that the GPU is done with the data that's 'owned' by the CPU side object - traditional destructors can be used of course by delaying the destruction on the CPU side texture object until the GPU is done with the data, but why keep the texture object around when only the pixel data is needed by the GPU, not the actual texture object?
A 'defer mechanism' is just the right sweet spot for a C like language IMHO.
OK? But now you need to… have two different mechanisms, when one can do both?
> In my mind, 'data' are the pixels in a texture (while the texture itself is an 'object'), or the vertices and indices in a 3D mesh (while the 'mesh' is an object), or the data that's read from or written to files (but not the 'file object' itself).
This “data” does not have a destructor, and thus is not a concern.
> C is all about data manipulation (the 'pixels', 'vertices' and 'indices'), less about managing 'objects'.
I think that would be news to every C program I’ve written.
> there are plenty cases in C++ where they are not
In which case you can just not have them, and not care.
> A 'defer mechanism' is just the right sweet spot for a C like language IMHO.
If by “a C like language” you mean “a language refuses to take any complexity off of the user’s back” then sure.
do {
SomeConnection *conn = conn_open(...);
if (!conn) break;
do {
SomeSession *sess = sess_open(conn, ...);
if (!sess) break;
do {
// ...
} while(false);
sess_close(sess);
} while(false);
conn_close(f);
} while(false);
The idea is to use "break" instead of "goto fail" because goto is bad.Why “BITITNT” instead of “BITINT”? I hope this is a typo, but “BITITNT” is repeated in the example code.
In my personal opinion, unless you're doing the project just for fun, it seems better to stick to C11/C17, at least for the next few years.
[1]: https://en.cppreference.com/w/c/23#C23_core_language_feature...
While in practice it may take long for some things to follow the standard, it is certainly a step in the right direction.
This is my humble opinion.
Where the compiler puts on stack the default value, whenever the function call doesn't include it.
Guess that won't compromise with anything.
int foo(int a, int b)
{
return a + b;
}
int main(void)
{
printf("Direct: %d\n", foo(1, 2));
int (*ptr)(int, int) = foo;
printf("Indirect: %d\n", ptr(1, 2));
return 0;
}
of course default arguments, if added, would have to be part of the function pointer type as well, making the above: int foo(int a = 1, int b = 2) // NOT REAL CODE, FANTASY SYNTAX
{
return a + b;
}
int main(void)
{
printf("Direct default: %d\n", foo());
int (*ptr)(int a = 1, int b = 2) = foo; // NOT REAL CODE, FANTASY SYNTAX
printf("Indirect default: %d\n", ptr());
return 0;
}
Unnamed function arguments would look silly (`int (ptr)(int = 1, int = 2)`?), but I would be radical then and only support default argument values for named arguments, probably.Edit: fixed a typo in the code, changed in-code comment.
So you could do foo({.bar=1}) and the rest would be default initialized.
struct bla_t {
int a = 23;
const char* hello = "Hello World!";
};
...and only allow comptime known constants for the default value declarations. struct bla_t {
int a;
const char* hello;
} blub = {
.a = 23,
.hello = "Hello World!",
}; struct bla_t {
int a = 23;
const char* hello = "Hello World!";
} blub = {
.a = 46,
};
The missing designated init item blub.hello would now be initialized to "Hello World!" by the compiler by looking up the default value in the struct declaration. Currently, missing designated init items are set to zero. And it would work in any other place where a bla_t is created: struct bla_t blob = {};
This would initialize blob to its default state of blob.a = 23 and blob.hello = "Hello World!".If I look at a piece of code that does
struct bla_t blub = {
.a = 46,
};
I know that everything except .a will be zero, without knowing anything about bla_t.With your proposal I would now get a 'helpful' default that might not be what I want at all.
Wait - no. What's wrong with
#define BLA_INIT_DEFAULT { .a = 23, .hello = "Hello World!", }
struct bla_t blub = BLA_INIT_DEFAULT;
blub.a = 42;
That's what your proposal would do under the hood anyway, just more explicit. #define BLA_DEFAULTS .a = 23, .hello = "Hello World!"
struct bla_t blub = {
BLA_DEFAULTS,
.a = 42,
};
...because C99 allows designated initializers to show up multiple times, but that's all a bit too much macro magic for my taste, I'd really prefer the defaults in the struct declaration. struct baz tmp = {.bar=1};
foo(&tmp);It would also free up zero as being an actual value instead of standing for 'default value'.
#define FOO_INIT_DEFAULT {.bar=1}
I don’t want some 'magic' struct that behaves different than all other structs on init.A static (default zero) struct lives in .bss, that entire section is zeroed on init.
If you want to have default values, it goes to .data where it will consume some ROM.
Also, once you initialize structs with values somewhere in the code (for instance with designated init), there's a high chance that the compiler will put a copy into the data section anyway, which is then memcpy'ed into the runtime struct (it depends on the compiler and compile options).
#include <stdio.h>
#include <stdlib.h>
struct params {
int a;
void *p;
};
void
f (struct params p)
{
printf ("p.a = %d, p.p = %p\n", p.a, p.p);
}
int
main (void)
{
f ((struct params){.a = 42});
f ((struct params){.p = f});
f ((struct params){});
exit (EXIT_SUCCESS);
}
Even though these are stack allocated, the missing fields are initialized to 0/NULL. The last case (no parameters) is new in C23. In C99 you had to use (struct params){0} to initialize it. https://en.cppreference.com/w/c/language/compound_literalSo if a dynamically linked library uses default values, and you make use of them, and the dynamically linked library decides to change its default values (e.g. a crypto library switches to more secure defaults), you don’t get the update until you recompile your own code.
That doesn’t sound great.
Isn’t it C# which uses this strategy, of embedding the defaults in the caller?
PS: I don’t think C is defined in terms of stack, and modern calling conventions use registers for at least the first few arguments.
char16_t s16[2 * sizeof mbs];
be
char16_t s16[sizeof mbs];
?
In Unix time, every day contains exactly 86400 seconds but leap seconds are accounted for.
It then provides an example for when the Unix time went from 915148800 to 915148800 after 1 atomic second on 1998-12-31T23:59:60.00So it's incorrect to say that Unix time does not include leap seconds.
C would have to change significantly to become a proper subset of C++, and pretty much no C program would still compile (because of int* c = malloc(sizeof int);).
https://softwareengineering.stackexchange.com/questions/1896...
No reserved words in PL/I: https://www.ibm.com/docs/en/epfz/5.3?topic=sbcs-identifiers
> Identifiers can be PL/I keywords or programmer-defined names. Because PL/I can determine from the context if an identifier is a keyword, you can use any identifier as a programmer-defined name. There are no reserved words in PL/I. However, using some keywords, for example, IF or THEN, as variable names might make a program needlessly hard to understand.
It would make more sense if C++ would reverse direction and become a superset of C (like ObjC choose to do from the start), but for C++ it's much too late now.
At least WG21 takes security more seriously.
I wish people took developer experience more seriously.
By the way, the C projects I used to work on 20 years ago, took an hour to compile, and it wasn't worse thanks to ClearMake sharing of object files.
In any case, even if templates make C++ slower to compile, I would rather have slower builds than more opportunities to keep security researchers busy.