Porting to GCC 14: C language issues
gcc.gnu.org
gcc.gnu.org
The issue with this set of compiler changes was that a lot of software builds successfully, but the new build suddenly has expected features missing. This happened with the previous attempt:
https://lists.fedoraproject.org/archives/list/devel@lists.fe...
That happened in 2016. In 2019, we undertook another attempt, this time with config.log/config.h diffing, but for various reasons stopped after fixing a few dozen packages, so it never went to the Fedora/upstream proposal stage.
After Xcode and Clang the changed the defaults, it was somewhat easier to convince people that we should get Fedora & upstreams ready for a future GCC change. It also helped that Gentoo has been pushing fixes for Clang compatibility upstream for a while (https://bugs.gentoo.org/408963, with a broader scope in https://bugs.gentoo.org/870412).
It's important to bear in mind the main change here is to stop wild C programmers from trying to run or release code that's almost certainly broken, because the diagnostic they got wasn't explicit enough. People can still force the old behaviour if they need to.
Heh. Surprisingly, I've seen some people claiming that it's impossible for a language without forward declarations to have a one-pass compiler that could handle calls of not-yet-defined functions, yet pre-ANSI C managed to do exactly that, in a pretty obvious (now that you know that it's in fact possible) way.
Granted, you can call that a second pass although that's not that different from emitting a function's epilogue IMHO.
This is what C compilers already do, in fact, to produce warnings when an implicit declaration doesn't match a later explicit declaration. But this is a best-effort warning only; it doesn't work if there is no declaration because the function is defined in a different translation unit, as I pointed out above.
That that information isn’t there is an implementation choice.
For example, they could have hacked it in like C++ did by mangling names (https://en.wikipedia.org/wiki/Name_mangling). That probably would have required supporting longer identifiers (IIRC archives limited them to 14 characters), but that’s doable.
You'd need to do this at link time instead, which would require completely overhauling the format of object files, dynamic libraries, static libraries so they carry information about the types of functions instead of just the symbol names. It's not an easy fix.
It does work in C ― that's what the include files are for (among other things), after all. So it's possible to be able to forward-declare functions outside the translation unit only (extern-declare?), those inside the translation units don't need to even if the compiler works in a single pass. And those external declaration could still be introduced at the very end of the translation unit and still would count. I dunno, seems like a pretty reasonable idea.
Part of it is backward compatibility: code written to assume 32-bit int could break. (Arguably such code is badly written, but breaking it would still be inconvenient.)
Another part is that C has a limited number of predefined integer types; char, short, int, long, long long (plus unsigned variants and signed char). If char is 8 bits and int is 64 bits, then you can't have both a 16-bit and a 32-bit integer type. Extended integer types (introduced in C99) could address this, but I don't know of any compilers that provide them.
What environment are you working in? Because I don't know a single half-recent compiler that does not provide stdint (uint8_t, ..., int64_t), but I mostly work with GCC/LLVM toolchains.
It's pretty common to develop part of an embedded C program under Linux or similar host environment. Better debuggers, better profiling tools, etc. And uint8_t and friends are particularly important when you're working cross-platform.
Sure, but isn't that just an implementation detail? Because I really don't care if my int64_t is internally typedef'd to "long long int" or "__m64", as long as there is a standardized interface to ask for it.
In that alternate universe, char could be 8 bits, short short 16, short 32, and int 64.
Only in flat address spaces, which excludes platforms like old 16-bit x86 or modern CHERI. There, "pointer difference within a single object" need not be the same size as "pointer reference".
In hindsight it would probably have been better to bite the bullet and make int 64 bits wide.
32-bit int is still arguably the native word size on x64. 32-bit is the fastest integer type there. 64-bit at the very least often consumes an extra prefix byte in the instruction. And that prefix is literally called an "extension" prefix... very much the opposite of native!
So could one make the argument that a 16-bit int ought to be the native word size on x64?
Personally, I find using int32_t in general to be an uglification of the code. I never use `long` in C code anymore, as it's sometimes 32 bits and sometimes 64 bits. I use `int` and `long long`.
Do I care about 16 bit code anymore? No. Very few programs would port between 16/32 these days anyway, no matter what the Standard says or how hard you try to write code portably.
Almost makes me want to add "typedef long long longer;" to some code that I don't intend anyone to maintain.
https://google.github.io/styleguide/cppguide.html#Integer_Ty...
I don’t love long long. As an amateur compiler writer, it hurts me. “long long” makes “long” both a modifier ( like unsigned is ) and a type. Yuck.
I wish it was i8, i16, i32, and i64 ( with u versions of each ). f32 and f64 for floats. Those are easy to understand and fairly easy on the eyes.
If those numbers are too noisy, the CIL ( .NET ) types could work. For example, i4 and i8 instead of i32 and i64. I do no love the look of i1 either though. I guess you could special case sbyte and byte as aliases.
Correct.
> fairly easy on the eyes
Not for me. It's a personal thing, I just don't like it. When I removed them all from my code, it was like I'd scraped the barnacles off my boat.
You would think casting something as signed or unsigned wouldn't promote to a int / unsigned int. Ditto for const and unconst.
A great battle ensued, and many champions were slain.
The value preserving folks carried the field, and the sign preserving folks changed their compilers.
Fortran is, as always, ahead of the game.
`-Dint=__INT64_TYPE__`
But with the migration to 64-bit machines, typically int stayed put at 32 bits.
I guess nobody wanted to introduce a new integral type between short and int; they had enough trouble dealing with code which assumed sizeof(long) == 4. I recall stumbling across a comment where the word "beint32_t" appeared where "belong" would have made sense in context..
It's almost as if they were not, in fact, intended for precise control of bitwidths in portable manner...
All in all, "it's very easy and straightforward to control the size and padding of a struct's fields in portable manner" is yet another C's imaginary advantage: it's not that straightforward or simple. The padding especially has always been a thorny issue.
If you have one `short` argument to your function maybe. But if you have a `short[]` array, you probably do care about the memory layout of that array. You might need it to be compatible with some particular data format that you're trying to read/write. Same with a field of a struct, if that struct is used for parsing. A lot of C code does parsing like this.
In fact, it works so well you don't even notice it working. If the C committee wants to improve C, they should make this an official feature.
P.S. Because of the lack of forward referencing, C code tends to be organized as leaves first, and the entry point at the end. This is simply backwards, the entry point should be at the beginning.
Someone shared with me their idea for a parallelized parser, where threads would parse different segments of the source file and then stitch their incomplete ASTs together.
I like the cut of your jib.
I've wondered why this hasn't been relaxed a bit when void* is used. Say you have these functions (modified from the link)
int compare (const char *a, const char *b) {
return strcmp(a, b);
}
int compare (const void *a1, const void *b1) {
const char *a = a1;
const char *b = b1;
return strcmp (a, b);
}
And then you have some FP of type int(*compare)(const void *a1, const void *b1) somewhere...If you have correct pointer const-ness, why isn't the first method the PREFERRED way of doing this? It's shorter, more clear, safer (calling compare directly and not through a FP still gets you proper type checking (you wouldn't call either of these example functions directly, but for others you might)), and IDE suggestions can better explain what the function is. I've thought several times about suggesting to compiler writers/the C committee to bless the first method. Is there some obscure hardware that the first would be incorrectly compiled or something (i.e. the calling convention for void* and foo* is different)?
And if the size of "void" and "char" aren't the same you cannot push two void* on the stack and pop two char*.
But, like I said: maybe I didn't understood what you said.
int compare_char (const char *a, const char *b) {
return strcmp(a, b);
}
int compare_void (const void *a1, const void *b1) {
const char *a = a1;
const char *b = b1;
return strcmp (a, b);
}
int (*call_void)(const void *a1, const void *b1) = &compare_char; //not valid in current C standard
Why can't the standard be changed so that it is valid to set call_void to &compare_char, so we don't need to write the longer compare_void? Are there architectures are still in active use where not all pointers are the same size in plain C (not worrying about C++) that would disallow this?AFAIK, x86, ARM, MIPS, SPARC, Alpha, and SuperH would all work fine calling compare_char though call_void. There could be some other issues, like near/far on 16-bit x86, but that would be orthogonal to implicit function pointer void* casting, and could still occur if call_void was set to compare_void. Do one of the more obscure embedded CPUs still in use have a varying pointer size?
A pointer to void shall have the same representation and alignment requirements as a pointer to a character type. [48]
<blah blah pointers to qualified vs unqualified, structs, and unions also have the same representation between themselves>.
Pointers to other types need not have the same representation or alignment requirements.
[48]: The same representation and alignment requirements are meant to imply interchangeability as arguments to functions, return values from functions, and members of unions.
The thing I've heard for systems where e.g. an int* and char* aren't equal is where the "int*" is the "primary" pointer type, and a char* has to add back the would-be-leading-bits or something. Though as per the above quote, char* vs void* would be safe? Doesn't help anything other than char* though.I think the most pragmatic solution is to just compile everything with no strict aliasing. You still get errors for accidental incompatible assignments but won't get bitten by the optimizer.
I'm curious how these dynamics play out in less popular compilers. If a compiler implemented VLAs in C99, they almost certainly still have that feature for backwards compatibility even after they support C11.
Is there any compiler which appeared on the scene between C11 and C23, and during that window, chose not to support VLAs and thus C99? It's not like C11 itself was very widely adopted, precisely because of the long implementation and industry rollout windows.
For these reasons, we propose to make variably-modified types
mandatory in C23. VLAs with automatic storage duration remain an optional language
feature due to their higher implementation overhead and security concerns on some
implementations (i.e. when allocated on the stack and not using stack probing).
from N2778 (https://www.open-std.org/jtc1/sc22/wg14/www/docs/n2778.pdf)> Variable length arrays with automatic storage duration are a conditional feature that implementations need not support
confused
void function(int n) {
int (*arr)[n] = malloc(sizeof(*arr));
}Um, C compilers already do that with arrays with compile-time lengths.
#include <stdio.h>
void main(void) {
char x[20][30];
printf("%zu\n%zu\n", sizeof(x), sizeof(x[0]));
}
prints 600
30
so you can have "char *y = malloc(sizeof(x)); memcpy(y, x, sizeof(x));" and it must work since C89 at least. The main problem with VLAs is that they make exact the stack frame size unknown until runtime which complicates function prologues/epilogues but that's the problem in the codegen part of the backend, the semantics machinery is mostly in the place already.P.S. And yes, uecker is a member of the ISO C WG14 and GCC contributor, according to his profile.
Of course, this runtime tracking has always been necessary for C99 VLA support, but I can easily see how it would be surprising for someone not deeply familiar with VLAs, especially given how the naive mental model of "T array[n]; is just syntax sugar for T *array = alloca(n * sizeof(T));" is sufficient for almost all of their uses in existing code. (In any case, it's obviously not "creating an object on the heap" that's the unusual part here!)
Well, is it much more machinery? IIRC doing
void f(size_t n) {
int x[n];
n += 1;
...
does not resize x, so there is no dataflow dependency or rather, x depends on a hidden const variable so no additional dataflow analysis necessary. void f(size_t n, int cond) {
if (cond) { n += 1; }
int x[n];
...
does resize x depending on the value of cond, so the size can't necessarily be known until the point where the type (int[n] in this case) is named.Also, the compiler has to make sure it keeps around implicit locals to store the variable layouts, so that code like
void f(size_t n) {
typedef int array[n];
n += 1;
array x;
...
functions as specified. This kind of pattern is one of the bigger things setting the feature apart from just "syntax sugar for alloca()".https://en.wikipedia.org/wiki/Variable-length_array#C99
If that phrasing isn't accurate I'm sure they'd appreciate an edit.
void example(int n) { printf("%zu", sizeof(int[n])); }
void example2(int n, int (*a)[n]) { }
This is optional: void example3(int n) { int x[n]; }I don't envy how much tech debt they have to deal with, while still trying to make it possible for new code to have a more modern experience. For all the excitement around new languages, we're still going to have a lot of C code for a long time, and difficult work like this should be appreciated.
>> So what is the current status? How many packages are going to be affected by this change? How are we going to track the progress?
> I see an unaudited rebuild failure rate of about 10% for rebuilds of source packages that produce arch-full binary packages. This number does not include packages which fail to build in rawhide without the compiler change. It includes packages which configure checks for something that we really do not support (like setproctitle or strlcpy). After the first pass completes, I'll have to do a second pass with the expected-to-be-undeclared functions gathered from the first pass filtered out. That should give us better numbers.
Sanity is such a rare thing these days.
This is how stable it is - they were still supporting C constructs from prior to C89 standardisation (which is when, I believe, void pointers were introduced).
"maybe consider using void * in more places (particularly for old programs that predate the introduction of void * into the C language). "
> The corrected standard C source code might look like this (still disregarding error handling and short writes):
>
> void
> write_string (int fd, const char \*s)
> {
> write (1, s, strlen (s));
> }
And disregarding the passed file descriptor! :)/s
void fun110 (char const * const *a) {}
char **a;
fun110(a);