When static makes your C code 10 times faster
mazzo.li
mazzo.li
This is a common issue with C code (including lots of code I've written). It's really easy to forget to const something, which forces the compiler to do global reasoning or to generate worse code. I've gotten into the habit of making things const unless I know I plan on mutating them, but I wish there was tooling that encouraged it. (BTW, this is something Rust does well by making things constant by default and requiring "mut" if it's mutable.)
EDIT: gcc seems to agree with me: you can see the optimized version here[1] and the unoptimzed version if you remove "const".
Isn't it valid to cast to non-const for a const, but only invalid to modify the const through the casted pointer?
Mostly seen in embedded space.
The core optimization is modulus % constant (and a power-of-2 as well). The static just enabled the optimizer to do better heavy lifting to get there. A const would've made the intent clear to the human reader and the compiler.
static const would've been best.
Advocating 'static' when your actual intent is 'const' does less experienced readers a disservice; they will assume that 'static' is meant to make things faster, and be disappointed when it doesn't work for non-constant values.
If anything, the title is meant to be read as "isn't it amusing that something apparently unrelated such as `static` causes a performance improvement".
That said, I have added a note clarifying this at top of the post now.
Global mutable state is pure evil. Don’t write globals.
It shouldn’t be easy to forget const on a global because a mutable global should produce immediate revulsion and nausea.
(I don’t really consider a const global to be “a global”. So ordinarily I’d just say globals are evil don’t write globals. But I’m trying to be explicit here.)
Nobody is being fooled about global state when you have a singleton database connection, event bus router, or network stack. I don't think your program is better when you pass in i/o functionality to every single class context in the constructor.
Similarly a mega-class that encapsulates everything your program does is also a code smell. There's no point to a private variable when everything can access it.
> I don't think your program is better when you pass in i/o functionality to every single class context in the constructor.
Abstracting over I/O transport is an excellent thing to do. This allows you to do things like easily record and replay a network stream. Which is useful for both debugging and automated tests.
I/O comes in a kazillion flavors. Networked, interprocess, serial port, file, synthetic, etc etc. It's definitely something that should be abstracted around and not doing so is something I've deeply regretted in the past.
> a mega-class that encapsulates everything your program does is also a code smell
Ok I agree it can have a foul odor. But even this can be advantageous.
Once upon a time Blizzard gave a GDC presentation about Overwatch. Kill-cam replays are notoriously difficult in video games.
Blizzard's solution to this was delightfully elegant. They made two copies of their world. One perpetually runs on latest. One takes snapshots of the world every N frames. When a player dies their viewport switches to the old snapshot which then simulates and renders for ~6-10 seconds. When the replay finishes or skips the viewport switches back to the main game, which never stopped receiving updates. This was a relatively trivial implementation given the complete lack of globals and singletons.
A mega-class lets you run parallel instances of your "world". It's also a nice pattern when you want to build-up and tear-down your world in-process and guarantee no stale state. For example when running tests you likely want certain tests to "start clean". It's nice to be able to do this without restarting the entire process.
I'll double-down that globals are evil. They are a sometimes necessary evil. Or the least bad choice. But my experience is that not using globals is almost always simpler, more elegant, more flexible, and ultimately preferable.
And if you are creating and scheduling all of your concurrency in user space, which is common for some types of server software, then passing what are effectively global objects down the call stack becomes a real mess and introduces a number of suboptimal behaviors in the code gen.
I’ve seen people try to design database engines, the high-performance kind that directly manage all the resources they use, that don’t use globals in a misguided attempt to adhere to this heuristic. The end result was a convoluted mess of indirection that just obscured the reality that all of those objects were mutable globals. If you care about performance then I/O isn’t very abstract; the code knows exactly what kind of device it is dealing with.
I don’t like mutable globals as a general rule, but for some types of software they are unambiguously the correct engineering choice and not using them would be a design defect.
There are definitely resources which are globally unique. But that does not necessarily follow that access should also be global.
Rust has some elegant patterns when working with embedded devices. For example GPIO pins could totally be stateful globals. But instead their passed around as types and Rust’s type system + borrow checker ensure correctness. It’s pretty neat.
I’ll assume you’re right for databases. My expertise is real-time VR video game type stuff. Which is also high-performance, but of a different variety.
I strongly agree that layers of abstraction compound into convoluted and inscrutable garbage. I loathe web development for this very reason.
In modern C++ it is straightforward to write wrappers in the style of unique_ptr that safely hide the DMA and life cycle mechanics, which means the average dev using them doesn’t need to know how it works, but someone has to write that code and it is necessarily global heavy because it references physical devices that have their own behavior in your address space. Under the hood, if a physical device is stomping on address space your code accesses, you need a way to both detect that an object is effectively owned by a particular DMA engine before touching it and immediately de-schedule the thread of execution until such a time as there is no concurrent DMA operation that might conflict with the code execution. This happens within a single thread, so no blocking or OS context switching.
A big part of database kernel internals is coordination and management of physical resources, which are global by nature.
LTO could handle this if you're compiling an executable, but not a library.
That's not the correct distinction, that's why I said default vs hidden visibility.
Libraries typically export more symbols but executables can also export them eg for plugins.
I worked on this as part of FreeBSD, which is why their base system is nowadays built with that flag enabled.
$ cat value.c
int value = 42;
int get_value() {
return value;
}
$ make value.o
gcc -c -o value.o value.c
$ nm value.o
0000000000000000 T get_value
U _GLOBAL_OFFSET_TABLE_
0000000000000000 D value
This symbol, not being `const` can be modified by any other compilation unit. $ cat main.c
#include <stdio.h>
int value;
int get_value();
int main() {
value = 123456789;
printf("%d\n", get_value());
}
$ make main.o
gcc -c -o main.o main.c
$ cc value.o main.o
$ ./a.out
123456789
Compiler when generating an object file has to assume the value of exported non-const symbol can change. It's necessary to tell the compiler that the value cannot change, either by not exporting the symbol by using `static` or making the value of it `const`. In example provided in your article `static` makes sense (or even `static const`) as I don't think there is a reason to export this global.The only option is when all TU are given at the same time to the compiler, or when LTO is used (in which case it is actually the linker doing the work).
Even then, this won't apply to libraries.
Exactly -- which is the case here. But implementing such a cross cutting implementation would probably be annoying, which is what I wanted to convey with
> I think they could concievably assume that the value of modulus won’t be changed in this case, since we’re producing an executable directly, but it’s probably annoying to have an optimization looking so far into the future of the compiler pipeline.
And arrays are bounds checked by default, no more issues with them (and strings are arrays)
let foo = "1234".to_string();
let mut foo = foo;
foo = "5678";
This only works if you have ownership of the variable though. And in practice it's a pretty good compromise, because it's pretty hard to do this by mistake.Note: const in rust is different to const in some other languages: it represents a value that is duplicated inline into each use site at compile time.
Static only means that the variable is local to the translation unit (the C file). The relevant difference in the example is actually the const-ness of the variable, which you may put explicitly, but which a powerful compiler can also infer here in the static case.
Other than this optimization possibility, const and static are orthogonal concepts. I'm not sure to what extent the article author is aware of this.
So the lesson should be: use const (or #define) if you mean to have a constant. It's still a good idea to also make things static, but the real reason for that is to avoid name collisions with variables in other C files.
const doesn't do this thanks to const_cast etc. maybe things have changed, but const on a file level variable doesn't do this reliably, or at least hasn't for considerable lengths of time.
of course 'static const' is the better answer. :)
From a technical standpoint, using const_cast to modify a constant is Undefined Behavior. This was explicitly specified that way to make constants inlineable by the compiler. Every compiler that I've ever used does this, it's a trivial but very effective optimization.
Proof by Godbolt: https://godbolt.org/z/djEdvee4s
Observe how GCC completely eliminates the contents of undefined_mutate_modulus (which it's allowed to -- UB means it can do anything with that function, and it chooses the simplest possible thing) rather than de-optimizing mod4_const like you suggested. Compilers are smart.
Having it local or const makes the compiler able to inline it and do a simple bitwise and with a constant.
So yes, make your variables static const by default (if you really need global).
Really wish it was the default.
I once had a gray beard chew me out over changing a function signature when revising things. I had to point out politely that it was static and that anyone who managed to link to the function had to be breaking a lot of rules to do so.
static doesn't imply the value can be known, only in some cases where its not modified in that compilation unit (which is trivial to detect with SSA)
const is something else.
Honestly if I as the programmer knew that I was really trying to select bits from a number I’d just use a binary and directly. In that specific situation I think the intent is more clear that way. Like:
//select the bottom 8 bits
unsigned bottom8 = val & 0xff;
When the compiler can assume your values don't change magically, it can optimize their use.
This is true for restricted pointers, for global-scope variables which can only be accessed in the same translation unit, for stuff in inlined functions (often), etc.
--------------------------------------------
const is a bit shifty. const makes the compiler restrict what it allows you to write, but it can still not really assume other functions don't break constness via casting:
void i_can_change_x_yeah_i_can_just_watch_me(const int* x)
{
*(int*) x = x + 1;
}
now, if the compiler sees the code, then fine (maybe), but when all you see is: void sly(const int* x);
You can't assume the value pointed to by x can change. See this on GodBolt: https://godbolt.org/z/fGEMj9Meoand it could well be the same for constants too. But somehow it isn't:
In your first example you have `int x = 1;` which isn't const, so the compiler has to assume that `f` may mutate it after const casting.
In your second example you have `const int x = 1;` which is const so the compiler can assume the value will never change.
Here are the references to linker scripts?
I believe this account of things is accurate (ignoring C++'s reference types for simplicity):
The C and C++ standards are written in such a way as to enable compilers to place constant data on read-only memory. Casting away constness is never illegal in and of itself, but if you do so and then assign to a variable which was declared as const, that is undefined behaviour.
Similarly, assigning into the character array of a string literal is undefined behaviour, whether or not const was used. Again that's to enable the compiler to make use of read-only memory.
Going in the other direction is safe, as preventing assignments doesn't introduce problems. That is to say, using a pointer-to-const type to point to a non-const variable poses no problem.
Related:
• https://stackoverflow.com/a/9079161/
• https://wiki.sei.cmu.edu/confluence/display/c/EXP05-C.+Do+no...
the situation is even worse with ELF dynamic libraries due to the interaction of two rules: a) by default, all functions are exported, and b) by default, all functions can be interposed, e.g. by LD_PRELOAD. here, if you specify -fPIC in the compilation arguments (as is required to produce a modern dynamic library), inlining is totally disabled. for small functions, the call overhead can be substantial.
I'd also ask him to make "switch" break by default.
Then I'd go kill Hitler or something.
Probably a module/namespace system would be the biggest improvement.
I made Firefox's code base able to compile with -Wimplicit-fallthrough. About one hundred fall through cases needed to be annotated and about 2-3 were actual bugs (though minor).
switch (foo) {
case 1:
case 2:
/* common body */
break;
case 3:
/* body 3 */
break;
}
It's not cognitively hard to make an empty case body fallthrough to the next run, but include an implicit break at the end of every nontrivial body.Supporting Duff's Device is not a compelling feature to support--irreducible loops are basically going to destroy any hope of optimization you might accrue.
which is probably the case in 90% of all my switch'es. Worst use of a time machine ever.
From Algol's linage point of view, a very good outcome.
Many languages also support something like
case 1,2,4:
or even case 1-10, 20-22, 34:
That further decreases the need for falling through.Also https://tvtropes.org/pmwiki/pmwiki.php/Main/HitlersTimeTrave...
A time traveler shouldn't have a limit on time to search, so this doesn't make sense as an objection to me.
That said, I'm more annoyed at C's design commitee to not being able to add new features such as having length-aware arrays in C in the language.
mod’ing by a power of two -- is equal to bitwise and of that number minus one!
All you need to do is keep the bits lower than that power of two, which is what the and will do."
the use of static is just a tool to inform the compiler that the value is a constant (which const /might/)
you can tell how much low-level optimisation they have done if they think its gonna change codegen reliably, or at ll.
Multiplication is also very fast, usually one or two cycles on larger chips.
I wondered if that's really still true, since I haven't done much assembly language programming since PowerPC was new.
Here, it says the M1 has 7-9 cycles latency for division instructions, but throughput of 2 cycles per.
https://dougallj.github.io/applecpu/firestorm-int.html
"The M1 is 10x faster than the Xeon at 64 bit divides. It’s…just wow."
So, given all of the other things that can slow you up, I wonder if it really makes sense to avoid division any more?
(I guess the energy efficient "Icestorm" cores have throughput equal to latency, so it's only the "Firestorm" ones where it's super fast)