Understanding GCC Builtins to Develop Better Tools [pdf]
manuelrigger.at
manuelrigger.at
Example:
{
__attribute__((cleanup(foo)))
int a;
…
/* foo(&a) will implicitly be called here */
}Some of the GCC builtin wrappers are outstanding, e.g. __builtin_memcpy on two-power sizes, since it can do things that wouldn't be possible otherwise.
Some GCC builtins, e.g. __builtin_va_start, can take away flexibility since it locks into the core codebase nontrivial code generation that's suboptimal in certain cases.
Most of the builtins are neutral, e.g. atomics api, which provides 42 different ways to say asm("mov"), 13 different ways to say asm("xchg"), and twenty different ways to asm("lock cmpxchg").
Other GCC builtins, e.g. __builtin_ia32_rdrand32_step, don't even have the multi-architecture benefit, only seem to exist because some folks disagree with the asm() keyword, and is naturally problematic from a security standpoint when the GCC devs get it wrong.
What's great about builtins like asm() is it's one simple general-purpose abstraction that puts the power of so many compiler internals into the hands of the user at once, rather than being another narrowly-focused case-by-case feature you need to beseech the core dev team to get.
I am not sure how your suggestion differs from how GCC is already being developed.
Depending on who you ask, this is a feature, not a defect. Inline assembly is your last line of defence against a compiler which thinks it knows better than you.
Clang/llvm go a step further and do parse and optimize your inline assembly; that makes me uncomfortable, but I’ve always subscribed to the “don’t touch user asm” mindset for compiler development.
Your main benefit of builtins over assembly is that no matter how hard I tried I couldn’t get the suggested replacement given by the parent for __atomic_compare_exchange to assemble on my arm64 server!
char DidIt; asm("lock cmpxchg" : "=@ccz"(DidIt), "+m"(IfThing), "+a"(IsEqualToMe) : "r"(ReplaceItWithMe) : "cc");
Two things complicate the AArch64 (Arm64) story for atomic builtins.
One is Arm’s weak memory model. This permits different, more efficient, implementations of compare exchange depending on which of the C++ ordering guarantees you need and ask for. A compiler will generate different code for the sequential consistent model than it will for acquire/release.
The other is that there are two ways to provide atomics on Arm, and which you have access to depends on the architecture revision implemented by your particular Arm hardware. If you have Armv8.1-A atomic operations available to you, you’ll use a similar interface to x86; a single instruction which looks like a cmpxchg. If you don’t have the atomic extension, you’ll instead be using a load-lock/store-conditional loop of ldrex/strex for all of your atomic operations.
Getting this exactly right is tricky, so where possible I’d encourage using the atomic builtins (or, of course, your language’s portable API, over assembly).
The biggest benefit is that if in future architecture technology moves forward again, the GCC and language built ins will let you move with it, while assembly leaves you relearning occasionally.
Usually what gets optimized the most, counter-intuitively, is the adjacent code that the compiler generates itself, rather than the handwritten assembly you supply (which can in fact be empty string). This basically gives you a means of controlling the compiler's internal algorithms for performing register allocation.
Take for example foo(). If that were written as asm() it'd imply asm("call foo" ::: rax, rcx, rdx, r8, r9, r10, r11, r12, xmm0, xmm1, etc.). If you know ahead of time that foo won't clobber those million different things, then explicitly using asm() can have a wild impact. I've seen just one of those alone cut down the code size of large functions by a third.