Some Obscure C Features
multun.net
multun.net
The printf specifiers:
%e Scientific notation 3.9265e+2
%a Hexadecimal floating point -0xc.90fep-2Often the numbers (e.g. coefficients of some degree 10 polynomial approximation of a special function) are not intended to be human-readable anyway. Such values were typically computed in binary in the first place, and the only place they ever need to be decimal is when written into the code, where they will be immediately reconverted to binary numbers for use.
https://randomascii.wordpress.com/2013/02/07/float-precision...
Say I have a static list of names and I would like to declare some struct type for each name. I also would like to create variables of these structs at some point, and I would always do so for the entire block of names. You could do something like this:
#define apply(fn) \
fn(name1) \
fn(name2) \
fn(name3) \
...
fn(nameN)
#define make_struct(name) struct name##_t { ... }
#define make_variable_of(name) name##_t name;
...
apply(make_struct); // This defines all the structs.
void some_function(...) {
apply(make_variable_of); // And this defines one variable of each type.
}
Yes, it is not pretty (it is the C preprocessor after all), but it can be very useful and clean.I used this trick to declare and generate code for an entire parser, it's lovely.
I recently switched to the other style of C-macro though:
#define X(A, B, C) ...
#include <math/apply_operators.h>
#undef X
This style is less powerful but doesn't require the annoying \ at the end of lines. enum foo {
FOO_THING_ONE,
FOO_THING_TWO,
FOO_THING_THREE,
...
FOO_THING_SEVEN_HUNDRED
};
// Using concatenate '##' and stringify '#' operators
#define FANCYCASE(X) case FOO_THING_##X: str=#X; break
const char *foo_to_str(enum foo myFoo)
{
char *str;
switch(myFoo)
{
FANCYCASE(ONE);
FANCYCASE(TWO);
FANCYCASE(THREE);
...
FANCYCASE(SEVEN_HUNDRED);
}
return str;
}Compile-time trees are possible without compound literals.
More than twenty years ago, I made a hyper-linked help screen system a GUI app whose content was all statically declared C structures with pointers to each other.
At file scope, you can make circular structures, thanks to tentative definitions, which can forward-declare the existence of a name, whose initializer can be given later.
Here is a compile-time circular list
struct foo { struct foo *prev, *next; };
struct foo n1, n2; /* C90 "tentative definition" */
struct foo circ_head = { &n2, &n1 };
struct foo n1 = { &circ_head, &n2 };
struct foo n2 = { &n2, &circ_head };
You can't do this purely declaratively in a block scope, because the tentative definition mechanism is lacking.About macros used for include headers, those can be evil. A few years ago I had this:
#include ALLOCA_H /* config system decides header name */
Broke on Musl. Why? ALLOCA_H expanded to <alloca.h>. But unlike a hard-coded #include <alloca.h>, this <alloca.h> is just a token sequence that is itself scanned for more macro replacements: it consists of the tokens {<}{alloca}{.}{h}{>}. The <stdlib.h> on Musl defines an alloca macro (an object-like one, not function-like such as #define alloca __builtin_alloca), and that got replaced inside <alloca.h>, resulting in a garbage header name. a = 0;
x = (type_t) { .y = a++, .x = a++ }
Now, figure it out the order of the field's assignment.We don't have to use mixed-up designated initializers to run aground. Just simply:
{ int a = 0;
int x[2] = { a++, a++ }; }
This was also new in C99 (not to mention the struct literal syntax); in C90, run-time values couldn't be used for initializing aggregate members, so the issue wouldn't arise.Code like:
{ int a = 0;
struct foo f = { a, a }; }
was supported in the GNU C dialect of C90 before C99 and of course is also a long-time C++ feature.https://gcc.gnu.org/onlinedocs/gcc/Initializers.html#Initial...
The doc doesn't say anything about order. I can't remember what C++ has to say about this.
extern foo n1, n2;
Is there any benefit of tentative definitions over this? struct args { int a; char *b; };
#define fn(...) (fn_((struct args){__VA_ARGS__}))
void fn_ (struct args args) {
printf ("a = %d, b = %s\n", args.a, args.b);
}
fn (0, "test");
fn (1); // called with b == NULL
fn (.b = "hello", .a = 2);
(As written this has a subtle catch that fn() passes undefined values, but you can get around that by adding an extra hidden struct field which is always zero).It was before C had built-in booleans and the author had defined their own, but true was:
void * true_ptr = &true_ptr;
true_ptr is a pointer to itself. So however many times you deference it: printf("%p\n", true_ptr);
printf("%p\n", &true_ptr);
printf("%p\n", *((void**)true_ptr));
printf("%p\n", *((void**)*((void**)true_ptr)));
You get the same pointer: 0x5555fefe715b
0x5555fefe715b
0x5555fefe715b
0x5555fefe715b
I still think that it's neat that, even with ASLR, you have an address at compile time that you know won't collide with address space of malloc results, or the address space of your stack.Also you can declare the pointer as const and the value it points to as const and, if your kernel faults on writing to readonly memory pages, you get a buggier version of a NULL pointer that only segfaults on write.
Also it takes a second to figure out why the position of the const matters even though the pointer's value is the value of the pointer, and why only one of these segfaults on write:
const void * const_pointer = &const_pointer;
void * const const_value = &const_value; cdecl> explain const void * const_pointer
declare const_pointer as pointer to const void
cdecl> explain void * const const_value
declare const_value as const pointer to void
cdecl>
The first must be the one that segfaults on write, IFF the compiler chooses to place it in the .text (as it should).Testing it on my machine with the following code seems to validate this hypothesis.
//file: test.c
#include <stdio.h>
const void * const_pointer = &const_pointer;
void * const const_value = &const_value;
int main()
{
printf("%p\n", const_pointer);
*(int*)const_pointer = 0;
printf("%p\n", const_pointer);
printf("---------------------------\n");
printf("%p\n", const_value);
*(int*)const_value = 0;
printf("%p\n", const_value);
return 0;
}
Result:
$ gcc test.c
test.c:4:30: warning: initialization discards ‘const’ qualifier from pointer target type [-Wdiscarded-qualifiers]
void * const const_value = &const_value;
^
$ ./a.out
0x55b29ebfc010
0x55b200000000
---------------------------
0x55b29ebfbdb8
Command terminated
As to why the first one doesn't also result in a segfault, I don't know.Edit: Actually I don't think the pointer is put anywhere, rather its value is stored in .data (non-read-only data) so it can be mutated without issue. Again though, my assembly isn't amazing.
In the second case the const_value variable itself is const-qualified and thus located in .rodata, but the pointer itself is not const-qualified so nothing prevents you from attempting to modify the data through that pointer. This is why you get a compiler warning about discarding the 'const' qualifier in the initialization. Since const_value is in .rodata, writing to it through the pointer causes a segfault.
As Sean1708 pointed out, it's more obvious what is going on if you place the 'const' qualifier immediately before the thing it's modifying, which is either the pointer operator or the variable name, never the type itself:
void const *const_pointer = &const_pointer;
void *const const_value = &const_value;
What would something like "const int" even mean on its own, anyway? There is no such thing as a mutable integer. It's the memory location holding the integer which may be either mutable or immutable. const void * const_pointer = &const_pointer;
void * const const_value = &const_value;
I've never really understood why people put const before the type, to me the following is far more obvious: void const* const_pointer = &const_pointer;
void* const const_value = &const_value;unsigned int x;
unsigned int * x_addr = &x;
unsigned int x_addr_addr = &(&x);
(or arbitrarily many levels of "address-of") and you'd just get the same address.
Excerpt from GCC man page:
Trigraph: ??( ??) ??< ??> ??= ??/ ??' ??! ??-
Replacement: [ ] { } # \ ^ | ~
Missing backslash on your keyboard? No problem, just type ??/ instead. if(condition)
printf("WTF??! value: %d", value);
...only to have the compiler nag me about it. That's pretty much the only situation I've come across trigraphs :)Fun fact: ISO 646 is also the reason that IRC allows these characters in nicknames. IRC was created in Finland, and the Finnish national standard had placed the letters Ä Ö Å ä ö å at those code points.
Edit: That doesn't explain the trigraphs for # ^ ~. I'm guessing some EBCDIC variants lacked those. Or some other computer vendor on the committee still supported some other legacy character set.
C++ code example on Godbolt: https://godbolt.org/z/ED6tXK
https://en.cppreference.com/w/cpp/language/operator_alternat...
For instance, the next line has a valid C construct:
return a, b, c;
This is particularly useful for setting variable when retuning after an error. if (ret = io(x))
return errno = 10, -1;
The possibilities are endless. Another example: if (x = 1, y = 2, x < 3)
...
But the comma operator really shines when used in conjunction with macros.And to be clear, even if it is common in JS, that still doesn't reduce the usefulness of the original comment, because we aren't talking about JS, we're talking about C, and the commonness of the features notes in the C ecosystem. There are plenty of obscure oft-ignored features on one language that are common in another. For example, taking someone's interesting C macro that allows some level of equivalence to functional map and apply and saying "that isn't very obscure, even lisp has that" is missing the point.
Surprisingly, you can override it in C++. I haven't seen anyone do it, but you can. If you find a good, productive override for the comma operator, please post about it.
>a[b] is literally equivalent to *(a + b).
Is this obscure? I thought that's pretty much the first thing you learn about arrays in C? It's pointers, all the way down.
foo[3]
*(foo + 3)
*(3 + foo)
3[foo]
They all do the same thing. foo[3] == (void *)((usize)foo + (usize)(3 * sizeof(*foo)))You can also subtract two pointers into the same array and get the distance (in elements, not bytes). a[b - a] == b.
typedef struct thing {
char a;
long long int nothing;
} thing_t;
#define P(x) printf("%s %c\n", #x, x)
int main(int argc, char **argv)
{
thing_t array[1024];
array[8].a = 'H';
P(array[8].a);
P((8[array]).a);
P((*(8 + array)).a);
}What do you mean? As far as I remember, C++ is very similar in this respect when it comes to array and pointer types.
Which is true, but I would argue an implementation that makes those operators behave differently is likely ill advised.
[1] Array designators, Preprocessor is a functional language and a[b] is a syntactic sugar.
void (*foo)() = 0;
void (*bar)() = (void *)0;
void (*baz)() = (void *)(void *)0; // Error
Can you guess why compilers reject the last line?C++ has stricter rules for null pointer constants, and thus only the first version is valid C++.
It's also not clear how to use that preprocessor example.
old but gold
Still makes me laugh
Here's what the example yields once preprocessed:
struct operator{
int priority;
const char *value;
};
struct operator operator_negate = {
.priority = 20,
.value = "!",
};
struct operator operator_different = {
.priority = 70,
.value = "!=",
};
struct operator operator_mod = {
.priority = 30,
.value = "%",
};The standard says:
"If the type of the operand is a variable length array type, the operand is evaluated; otherwise, the operand is not evaluated and the result is an integer constant."
But what does it mean to evaluate the operand?
If the operand is the name of an object:
int vla[n];
sizeof n;
what does it mean to evaluate `n`? Logically, evaluating it should access the values of its elements (since there's no array-to-pointer conversion in this context), but that's obviously not what was intended.And what about this:
sizeof (int[n])
What does it mean to "evaluate" a type name?It's not much of a problem in practice, but it's difficult to come up with a consistent interpretation of the wording in the standard.
When you say "n", syntactically, you have a primary expression that is an identifier. So you follow the rules for evaluating an identifier, which will produce the value. C doesn't describe it very well, but the value of the expression is the value of the object. In terms of how it is implemented in actual compilers, this would mean issuing a load of the memory location, which is dead unless `n` is a volatile variable.
int n = 42;
int vla[n];
sizeof vla;
not `sizeof n`). (It doesn't look like I can edit a comment.)Logically, evaluating the expression `vla` would mean reading the contents of the array object, which means reading the value of each of its elements. But there's clearly no need to do that to determine its size -- and if you actually did that, you'd have undefined behavior since the elements are uninitialized. (There are very few cases where the value of an array object is evaluated, since in most cases an array expression is implicitly converted to a pointer expression.)
In fact the declaration `int vla[n];` will cause the compiler to create an anonymous object, associated with the array type, initialized to `n` or `n * sizeof (int)`. Evaluating `sizeof vla` only requires reading that anonymous object, not reading the array object. The problem is that the standard doesn't express this clearly or correctly.
"If the size expression of a VLA has side effects, they are guaranteed to be produced except when it is a part of a sizeof expression whose result doesn't depend on it." (from: https://en.cppreference.com/w/c/language/array)
In the example this is not a problem, because int[printf()] means you must call printf() to get the return value and determine the size of the array.
int b[const 42][24][*]
but you can get more fun with variably modified array parameters instead of just constants like 42. For example, double sum_a_weird_shaped_matrix(int n, double array[n][3*n]) {
double total = 0;
for (int x = 0; x < n; ++x) {
for (int y = 0; y < 3 * n; ++y) {
total += array[x][y];
}
}
return total;
}
has a variable and a more complicated expression in those positions.But those variably modified parameters can have arbitrary expressions in them, like
int last(size_t len, int array[restrict static (printf("getting the last element of an array of %zu ints\n", len), len--)]) {
return array[len];
}
C++ denies us this particular joy which could have made function overload resolution even more fun.So if you do
void func(int x[10]);
You're free to call it like int k[5];
func(k);
And you won't get any warnings. Unsettling! void func(int x[static 10]);
must be called with an argument that is a pointer to the start of a big enough array of int. I can't get recent GCC or Clang to warn on violations of this, though. void foo(int *p)
{
func(p);
}
How can the compiler know if `p` points to space for 10 integers?The C FAQ is pretty old though, I’ve always wondered how much of that advice changed in C99/C11... from cursory googling things don’t seem to have changed much.
struct foo { char a[5]; };
void f(struct foo x) { x.a[4] = '\0'; printf("%s", x.a); }
int main(void) {
struct foo x;
memcpy(x.a, "too big", sizeof(x.a));
f(x)
printf("%s", x.a); /* read past end of x, crash */
return 0;
}
I'm thankful they didn't make structs decay into pointers!I needed to compare and older and newer version of some file from the RCS, so I saved temporary copies named "new" and "old". diff told me what I needed to know, but I failed to delete those temp files.
Hours later I typed "make" to build my program and got all sorts of errors deeply nested in some library function. Did someone misconfigure the server I was on? OK, maybe it is an incremental build problem? etc. It took took long to figure out the problem.
It turns out that during compilation, as one of the library .h files was being scanned, it contained #include <new>, which picked up the junk file in my working directory instead of using the C library.
Does anyone have any good use cases for it?
https://elixir.bootlin.com/linux/latest/source/drivers/clk/s...
Though the compile-time magic with structs and functional macros are so tempting that I feel like it's high time to do some C.
I used this feature recently. I had several arrays of the same size and type, and the size was determined at runtime. The VLA typedef let me avoid duplicate type signatures which I find more readable.
int N = atoi(argv[1]);
typedef int grid[N][N];
grid board;
grid best;
grid cache; int x = 10;
while (x --> 0) {
printf("%d ", x);
}"!=" would be the inequality operator... :-)
Wow. This has to be the best C obscurity that I've ever seen.
Notice how the list is sorted from actually obscure to less interesting.
I even took the time to write a pseudo disclaimer above this one :D
type quaternion float[4];
quaternion SLERP = {1.0, 0, 0, 0};
typedef float quaternion[4];
?That's not a VLA.
int main()
{
int a = 8;
{
int a = 4;
/* a is only scoped to this block */
}
printf("%d", a); /* prints 8 */
}
It is also why C++ is not a strict superset of CCan you explain? That code in C++ also scopes ‘a’ to the block.
EDIT: I see you’ve edited the code, but I think it’s still true in C++. I’ve often done that for RAII and unless I’m mistaken it works just as well when shadowing variables like you’re doing as when not.
A perhaps obscure feature is that you can "unshadow" a global variable like this:
#include <stdio.h>
int global = 0;
int main()
{
int global = 1;
{
extern int global;
printf("%d\n", global); // prints 0
}
} #include <stdio.h>
int main() {
printf("%d\n", sizeof('a'));
}
It's the second thing C++ mentions in its list of incompatibility with C (the first is "new keywords").A more obscure difference is this program:
int i;
int i;
int main() {
return i;
}
It's legal C but not legal C++. int main() {
return 4//**/2
;
}So if sizeof ('a') == 1, then either you're compiling as C++ or you're compiling as C under a rather odd implementation.
Both POSIX and Windows require CHAR_BIT==8. The only systems I know of with CHAR_BIT>8 are for digital signal processors (DSPs).
If you want to tell whether you're compiling C or C++:
#ifdef __cplusplus
...
#else
...
#fi