Lesser known tricks, quirks and features of C
jorengarenar.github.io
jorengarenar.github.io
long result;
register long nr __asm__("rax") = __NR_close;
register long rfd __asm__("rdi") = fd;
__asm__ volatile ("syscall\n" : "=a"(result) : "r"(rfd), "0"(nr) : "rcx", "r11", "memory", "cc");
The above is basically how you might implement the close(int) syscall on x86-64.You don't need to be doing embedded programming to find it useful to dip down into assembly like that (though syscalls are perhaps a bad example, even for a syscall not provided by your C library -- that library probably provides a `syscall` function/macro that does all this in a platform agnostic way).
Also, "%.*" is extremely useful with strings, i.e., "%.*s". Your code base should be using length delimited strings (basically `struct S { int len; char* str; };`) throughout, in which case you can do `printf("%.*s\n", s.len, s.str);`
The only good use of register these days is for project-wide globals in embedded contexts. IIRC one example of this is the decompilation of Mario 64, where a certain register ALWAYS contains the floating point value 1.0.
[Edit: Although, on second glance, from https://gcc.gnu.org/onlinedocs/gcc/Extended-Asm.html#Input-O...:
> If you must use a specific register, but your Machine Constraints do not provide sufficient control to select the specific register you want, local register variables may provide a solution (see Specifying Registers for Local Variables).
indicates in the r10 case maybe you _must_ use the syntax I gave?]
My preference is for the syntax that requires looking up fewer tables in GCC docs, but as I said, the version you prefer is fine too.
If your point is that I don't use result, that's because it's a snippet written into Hacker News. I didn't write the code to convert it into an errno and return -1 on error, etc., but doing so would be perfectly valid, and safe from your reasonable optimizer concerns.
register int *p1 asm ("r0") = …;
register int *p2 asm ("r1") = …;
register int *result asm ("r0");
asm ("sysint" : "=r" (result) : "0" (p1), "r" (p2));
and int t1 = …;
register int *p1 asm ("r0") = …;
register int *p2 asm ("r1") = t1;
register int *result asm ("r0");
asm ("sysint" : "=r" (result) : "0" (p1), "r" (p2));
In my example, nr is rax and listed in the input section, rfd is rdi and also listed in the input section, and result is rax and listed in the output section (I even used your preferred syntax for specifying rax here). Using result after the syscall asm statement is perfectly valid. asm {
mov RAX,__NR_close;
mov RDI, fd;
call syscall;
}
The compiler automatically keeps track of which registers were modified.Small note: your code seems to be calling a wrapper function rather than running the syscall instruction directly so it could probably be even more succinct. Is this valid D?
asm {
mov RAX, __NR_close;
mov RDI, fd;
syscall;
}https://www.felixcloutier.com/x86/syscall
> Why can't GCC and clang have this?
I don't know. My brain bleeds out my ears every time I have to deal with that.
Because UNIX culture sucks at inline Assembly.
D's approach was widely used across PC, Amiga and Atari ST compilers.
Besides D, that is what Delphi, C++ Builder and Visual C++ still use nowadays, for the targets that still support inline Assembly. For 64 bit targets, they use intrinsics, which are still preferable to UNIX's approach.
There are tricky contingent reasons, like the fact that GCC proper doesn’t try and actually isn’t even able to parse assembly—and it would need to here in order to know that it can’t put fd in RAX or in any location whose address computation includes RAX.
(GCC’s codegen in general works by stitching together fragments of assembly as text and feeding them into the system assembler. Which is nice and UNIXy and all, but I’m not sure much is won by making the compiler agnostic to the actual bit patterns of the instructions when these days it is still aware of almost every other facet of the instruction set and quite a few microarchitectural details besides.)
But the fact of the matter is that GCC’s approach also permits better optimization, by telling the register allocator exactly which registers are clobbered by the inline assembly snippet and where the computations of the input should aim to place their results. You frequently end up with no MOVs at all, although Clang is better at this part than GCC.
(Do not be fooled into thinking that an automatic "register asm" variable is pinned to the specified register for all of its scope—it’s only guaranteed to be placed there at the start of any inline assembly blocks, the register allocator can move it out of the way in between them.)
I think the x86 code doesn’t actually need the "register asm" syntax at all, FWIW—you should be able to just add "a"(__NR_close), "d"(fd) to the inputs. GCC on RISC archs doesn’t usually define a separate constraint for each individual register, though, so on those it can be necessary.
#include <sys/syscall.h> /* for __NR_exit */
int main(void)
{
long ret;
register long nr __asm__("rax") = __NR_exit;
register long ec __asm__("rdi") = 99;
__asm__ volatile("syscall\n" : "=a"(ret) : "r"(ec), "0"(nr) : "r11", "rcx", "cc", "memory");
return 0;
}
Build: clang -s -O0 -o ./t ./t.cDisassemble to double check: objdump -M intel -D ./t
Run: ./t || echo $?
Output: 99
The docs you link basically say that the special global "1.0" register mentioned in https://news.ycombinator.com/item?id=36551728 wouldn't be supported in clang -- so much for the only good usage!
long long int x;
int long long x;
long int long x;
typedef int myint;
int typedef myint;
const char *s;
char const *s;
const char * const volatile restrict *ss;
const char * volatile const restrict *ss; char *something; /* no null termination */
size_t something_length;
printf("%.*s", (int)something_length, something);
Unfortunately, the .* argument has type (int), not size_t, and it's signed… but if that's not a problem this is a great way to format non-\0-terminated strings.(And of course you can also use it to print a substring without copying it around first.)
In particular there are some APIs where you need to submit the whole format/string in one go, an example from libc is syslog(), but other libraries do this as well occasionally.
[0]: https://nebelwelt.net/publications/files/15SEC.pdf
Excellent.
typedef union c2_rect_t {
struct {
c2_pt_t tl, br;
};
c2_coord_t v[4];
struct {
c2_coord_t l,t,r,b;
};
} c2_rect_t; if (error_condition)
return *result=0, ERR_CODE;
So, back to writing lots of statements.Sometimes you can still avoid multiple statements by rearranging your code otherwise. For example, in your case, you can set *result=0 at the beginning of the function. Other times, you can also cram the assignment inside the condition using short circuit evaluation; this trick somehow seems more palatable to normies than the comma operator.
return errno=ENOENT, (FILE*)0;
I don't know if anyone uses this style.In K&R, they say that it's mostly used in for loops, such as
for (i = 0, j = strlen(s) - 1; i < j; i++, j--) ...
So that is where I use it now. return scoped_lock{foo_mux_}, return foo_;
Also nobody likes it. return scoped_lock{foo_mux_}, foo_;But that link no longer works.
I asked on SO why C characters use `'` on both ends instead of just one (e.g. why not just `'a` instead of `'a'`?). This seems to have been the biggest reason.
Multi-character constants are historical baggage.
You have to parse them the same way in a character literal as in a string literal, anyway.
I just figured out that
1. `\0` and octal numbers share the same prefix
2. Octal numbers can have 1-3 digits (not fixed)
So maybe it’s more tricky than I thought.
'\0' is just another octal escape sequence, not a special-case syntax for the null character.
"\0", "\00", and "\000" all represent the same value; "\0000" is a string literal containing a null character followed by the digit '0'.
Hexadecimal escape sequences can be arbitrarily long. If you need a string containing character 0x12 followed by the digit '3', you can write "\x12" "3".
It is more that it is an implementation detail of some compilers that was then (ab)used by certain platforms.
C11 added _Generic to language, but turns out metaprogramming by inhumanely abusing the preporcessor is possible even in pure C99: meet Metalang99 library.
I'm actually working on a library doing just that! It's still in very (very) early development, but maybe someone may find it to be interesting. [1]Link [2] is the implementation of a vector. Link [3] is a test file implementing a vector of strings.
[1]: https://github.com/jenspots/libwheel
[2]: https://github.com/jenspots/libwheel/blob/main/include/wheel...
[3]: https://github.com/jenspots/libwheel/blob/main/tests/impl/st...
https://abissell.com/2014/01/16/c11s-_generic-keyword-macro-...
I looked through this list, and I gotta ask, which items exactly do you find scary? Most other popular languages have similar, if not worse, quirks than the ones in this particular list.
* the comma operator (why not just allow blocks to be used as expressions?)
* multi-character constants (yay implementation-defined behavier)
* interlacing syntactic constructs (why?????)
* the fact that the --> operator works
* the idx[arr] thing
* the fact that you need a dirty hack with enums to check something at compile time
* flat initializer lists (why?)
* void pointers (instead of a proper type system)
* the syntax for function types
* the fact that X-Macros are needed
Isn't all of reference-implementation language like Rust and Python all implementation-defined?
> * the fact that the --> operator works
I pointed out elsewhere (with a link) that that specific construct works in most popular languages, like Java.
> * the fact that you need a dirty hack with enums to check something at compile time
It's not a check, its an assert. The alternative, for languages like Java and C#, is not having compile-time asserts at all.
> * the fact that X-Macros are needed
They are not needed, and indeed, equivalent functionality isn't in most popular languages (or anything but, Go, I think).
Since many of your complaints are applicable to other languages, it seems, to me anyway, that what you know of C is what you've read online in popular forums.
Languages like Java literally copied most of their syntax from C, so obviously they have the same problem and I hate them too.
Ditto for other languages that don't have compile-time asserts.
X-macros are a poor substitute for… well, actual macros. Or just metaprogramming in general.
The '-->' nonoperator.
Defining structs in a function declaration gives me the creeps. Same goes for the nested struct definition interpretation like come on. Void pointers in general stress me out. That's one I knew about before reading this. Makes bookkeeping hard.
I get that C is basically sugar around ASM. I have respect for it. At the same time, it feels so hazardous to write. Everytime I think I know enough I stumble on some weird UB landmine, or find some bizarre library that despite understanding 99% of the code the remaining 1% conceals everything I care about.
The quirks in many other languages can be rough, but I feel like Cs leave developers in a scarier place given the risk factors
it's the same in most languages. Works identically in Java, for example: https://www.jdoodle.com/iembed/v0/JL5
See?
> Defining structs in a function declaration gives me the creeps.
Both Kotlin and Go allow defining classes anywhere - they're anonymous.
> Same goes for the nested struct definition interpretation like come on.
Come on ... what? That's not a footgun. Because C is strongly typed, anytime you do a nested definition and expect it to be in the scope of the parent struct, you'll get an error if you try to create a new struct with the same name.
> Void pointers in general stress me out.
Fair enough, casting to and from void pointers explicitly throws away the type information.
> The quirks in many other languages can be rough, but I feel like Cs leave developers in a scarier place given the risk factors
The two quirks you complained about here are in other languages.
While C has problems, there is not much in this list that is problematic.
Although even things in the C FAQ get overlooked and forgotten.
eg: a 'Day of the Week' one liner such as
dow = (y + y/4 - y/100 + y/400 + m["-bed=pen+mad."] + d) % 7;
with int day (d), month (m), year (y) inputs and single int day of week output has a slight twist that causes a few head scratches for some.
y -= m < 3; dow = as above, etc;
I was focused on an 'obscure' yet fundemental part of the C language pointer | array equivalence.
Addendum: https://c-faq.com/misc/zeller.html
Constructs fully determined at compile time have some benefits. But const in C is weaker than constexpr that C++ has.
As prog-fh summarizes on https://stackoverflow.com/questions/66144082/why-does-gcc-cl...
> "The const means « I swear that I won't change the value of this variable » (or the compiler will remind me!). But it does not mean that it could not change by another mean I don't see in this compilation unit; thus this is not exactly a constant."
One example of invalid C:
const int mysize = 2; const int myarray[mysize];
gcc: error: variably modified ‘myarray’ at file scope
clang: warning: variable length array folded to constant array as an extension [-Wgnu-folding-constant] const int myarray[mysize];
* Good news: C can do compile time constant structs and array with deep self-references.
Yes, in C you can define and fully declare complex data structures that are accepted as compile-time constants, including pointers to parts of itself.
See "self-contained, statically allocated, totally const data structure with backward and forward references (pointers)?" for a previous example at https://stackoverflow.com/questions/47037701/can-c-syntax-de...
-----------------
I used this for a game on a retro machine where such a data structure avoids code which would have been several times (perhaps 10 times) bigger: https://github.com/cpcitor/color-flood-for-amstrad-cpc/blob/...
Here's another take showing two variants: where overall construct is an array then a struct: https://gist.github.com/fidergo-stephane-gourichon/792c194e1...
https://news.ycombinator.com/item?id=23445546 - Tic-Tac-Toe in a single call to printf
I wonder why developers tend to be so self deprecating
Just like being a sales person doesn't automatically make you overconfident, but being overconfident makes you a good sales person.
> ((struct Foo){}).x = 4;
Do such lvalues have any real use?