The C++20 Naughty and Nice List for Game Devs
jeremyong.com
jeremyong.com
Ok, I was also surprised to see co-routines on the nice list, but I don't have direct experience there. I normally see complaints about them. I would like them to be good because some code is easier to express that way.
The author talks about the code bloat, beacause of "an API that encourages custom formatter specification to live in a template". But at the end he mentions the standard solution to this problem:
> A preferable interface (I use, but also others AFAIK) is to check the type in a template (no choice there), and dispatch the formatting routine to somewhere that lives in a single translation unit.
So what prevents you from doing this with <format>? As I understand, the implementations of parse() and format() of std::formatter don't depend on the template parameters and can delegate to non-template functions residing in one CPP file. You can also provide additional wformat_parse_context/wformat_context overloads if you need wchar_t support.
The alternatives are worse: un-type-checked printf or the horrible stream interpreter system (std::cout << “foo”) which was a cute but bad idea in 1985.
Don't get me wrong, the fmt library is very nice, but you can't deny its effect on compile times.
In my corner of the C++ world though, I am so, so excited for <format> in 6 years or however long it will take us to move to C++20.
https://www.reddit.com/r/cpp/comments/o94gvz/what_happened_w...
Since I rarely compile all my code at once (usually just a single file followed by a re-link) compile time doesn’t matter much. And that’s even though while editing or writing code I don’t have all the slowdown bloat of an IDE so compile time is more noticeable.
* https://vitaut.net/posts/2020/fast-int-to-string-revisited/ * https://github.com/fmtlib/fmt/pull/1882 * https://vitaut.net/posts/2020/optimal-file-buffer-size/
Here's just one example: https://aras-p.info/blog/2022/02/03/Speeding-up-Blender-.obj...
Is that really 'an average' for modern AAA game?
Damn. That's an order of magnitude bugger then I'd imagine
Just C++ 11,375,669
Total (of everything) 31,379,114
That’s fairly representative of just the tooling side of things for a AAA engine. That’s not counting the logic of the game itself.
a big part of working on them as a generalist ends up being the ability to know how to even navigate something like that (especially since they're often haphazardly documented)
(part of it is that most of the games "fork" the engine rather than using it as a standalone thing)
it's probably not everyone on the team building that whole thing each time, but yea. hundreds of solutions and millions of LOC isn't unusual
*i just did a quick check with unreal's source, it's ~20 million LoC (assuming I didn't mess up the filtering somehow)
Large code bases built and linked with open source toolchains have solved this with, for example, thinlto. And by "large" I mean orders of magnitude larger than the mentioned game.
But yeah, outside of some of the largest of FAANG software, there aren't a lot of codebases that can truly justify 100m lines of code. I'm sure GTA VI is over 100m lines but can be cut down to 10m if that was a priority (it never is)
Especially for games or OS development, you might have shifting toolchains and SDks. Different teams may move out of sync because different teams want different things at a given time.
While I wouldn't want to work on a 10mloc C++ code base either, it sounds totally realistic to me.
Also, sometimes you need to iterate on some very core .h file and touching any of those brings the whole house of cards down and triggers a full or nearly full rebuild.
BTW how about std::is_constant_evaluated()? I assumed it would help folks who do heavy physics simulations, but looks like not listed in the article.
Of course now that std::bit_cast exists it's the safe thing to do (but then there's still C code that's compiled in C++ mode which was even recommended by Microsoft because the Visual Studio team couldn't be bothered to keep their C compiler in shape until a little while ago).
GCC maintainers: Hold my Jolt
The problem isn’t that compilers won’t implement the feature (that would take more work); the problem is that it’s processor-specific.
The spec doesn’t mandate many specific bit-ordering layouts (some are, such two’s complement representation being mandated which was just added, &obj == &base, I think nullptr has to be 0, etc) rather than trying to make everything a PDP-11.
I hate this about C++! In C you can initialize them in any order, and this allowed us to write nbdkit plugins in a very natural way:
static struct nbdkit_plugin plugin = {
.name = "myplugin",
.open = myplugin_open,
.get_size = myplugin_get_size,
.pread = myplugin_pread,
.pwrite = myplugin_pwrite,
/* etc */
};
where the order is not related to the order the fields appear in the struct (that has to be maintained for ABI reasons), and not all fields need to appear (the others are initialized with 0/NULL).For C++ we have to do this mess:
https://bugzilla.redhat.com/show_bug.cgi?id=1418328#c3
Anyway my question is .. why is this, C++ people?
...no chaining of designated initializers:
const bla_t bla = { .a.b.c = 23 };
...and no array indexing: const blub_t blub = {
.arr = {
[4] = 23,
[2] = 1
}
};
...all those limitations taken together, and the C++ designated initialization feature is pretty much useless except for the most trivial structs - while in C99, designated initialization really shines with complex, nested structs.The funny thing is that none of those limitations would be required. Clang had supported full C99 designated init in C++ mode just fine for many years before C++20 appeared.
We find it pretty useful even with the limitations. Certainly not “pretty much useless.”
I think you can do a similar sort of array initialization in C#, but definitely not chained initializers. Those are both so useful, but I can see why they aren't included in C++/++
struct foo {
int a = 0;
int b = a+1;
}
If the compiler just did the initialization in the order of declaration, regardless of the order in the initialization list this would not do what you expect:
struct obj {
int a;
int b;
} int ival = 0;
auto o = obj {.b = ++ival, .a = ival};
o.a would not equal o.b.I would like to have the initialization syntax of C because then one could reorder elements (say for packing reasons) and the designated initialization would “just work”…except it wouldn’t.
C++ designated initialization does buy you two things: 1- documentation, but more importantly 2- if you do reorder a struct or class data members the compiler will warn you that your initialization lists are now invalid rather than silently failing. I don’t know how to even find them all in a large code base any other way!
[0] https://gitlab.com/nbdkit/nbdkit/-/blob/cd761c9bf770b23f678f...
Even though many only know it with C, Fortran, Pascal and Ada were also common when C++ was gaining adoption.
It was never about being a polyglot compiler.
I also thought that the behaviour as standardized was useless, but recently I started writing more minimalist code eschewing constructors where aggregate initialisation would suffice, and I haven't really missed the ability to reorder initializers or skip them.
From what I can tell, the snippet you posted would compile fine in C++20 mode.
struct S {
int A;
int B = A;
int C;
};
int i = 0;
S s = {.C = i++, .A = ++i};
What would you expect code like this to do?This would still be a lot better for 99% of real world use cases than requiring the programmer to manually place the items in declaration order.
f(a++, b++)
That footgun doesn't seem to have very much impact in the real world that I've seen. Largely because people do not write complex expressions like that anymore. Since we already mostly avoid such expressions, we may as well take the benefit for designated initializers.It is a much smaller footgun though and I don't recall ever being bitten by it. I have definitely been bitten by member initialisation order though multiple times. Generally because of questionable designs where one member is passed as a parameter of another member and it hasn't actually been initialised yet.
Really I think the answer is to initialise in the order that initialisation is written (this is what Rust does) rather than declaration order. But that would be a breaking change so I guess they opted for the conservative choice.
On the other hand evaluation order of aggregate initializers is well defined (left to right) and there was little appetite to weaken it.
It might still happen in the future of course.
The problem is, someone thought this was a good idea, but the act of supporting it ruled out a lot of more-useful future improvements.
> Personally, I find code that leverages ranges harder to read, not easier, because lambdas inlined in functions introduce new scopes that have a strong non-linearizing effect on the code. This isn’t a criticism of ranges per se, but certainly is a stylistic preference.
Does anyone know what “non-linearizing” means here?
It can especially create problems when the lambda captures a variable by reference which gets mutated and/or deallocated before the lambda runs, and the developer didn’t plan for mutation or deallocation.
Or (a problem with lambdas, but not “non-linearizing”), if the lambda captures a variable by value (copies the value) and mutates it, and the developer expected the mutation to persist outside the lambda.
But the sort answer is all the other operators are automatically generated from that one if it is defined. So it makes the code simpler. And for many types <=> isn't much more complicated than the others
I think that the dramatic consequences are only understandable if you succumb to mimetic contagion.
The consequences are real but not dramatic and possibly not even measurable in many workloads.
It just means that you’ll have an extra sign extension (one of the cheapest ops the CPU has) in a subset of your loops, namely the ones that had a 32 bit signed induction variable and the compiler could reason about that variable but only if it also could assume no wrapping. That’s a lot of caveats.
Most loops will be unaffected by making signed integer overflow defined. Anything that’s not in a loop will almost certainly be unaffected by this change. If you use size_t as your indices then you’ll definitely be unaffected.
So yeah. “Dramatic consequences”. I wish folks stopped exaggerating. There’s nothing dramatic here. It’s a fraction of a percent of perf maybe.
(Amateur C programmer silly question) I think I understand it as if we increment the variable (i+10) and use it in an if condition. With UB the compiler could skip that code altogether and assume it will never be reached?
It’s more like this. If you say A[I] where I is 32 bit signed and you’re on a 64 bit target, then this lowers to:
- sign extend I to get a 64-bit value
- multiply it by the size of A’s element type
- add that to A
- then do the access
The last three steps will be just one instruction in the common case on arm and x86. The first step will require a separate instruction on x86.
The compiler can kill the sign extend if it’s sure that the integer value cannot be negative. That’s hard to prove. But you can almost prove it if you see code like:
for (int i = 0; something; ++i)
It looks like i starts out as zero and only grows! So it has to be positive! So if you say A[i] then no sign extend needed!
But wait, what if ++i overflows?
With signed int UB, the compiler can just assume it won’t overflow. And then it can prove that i is nonnegative. And then it can kill the sign extend on those CPUs where it’s not free, like x86.
I’m a compiler writer. I know how valuable this optimization is. Namely, it’s the tiniest of benefits on some program/CPU combos. Modern languages like Java or Swift just give this well defined semantics and call it a day because this isn’t a good hill to die on. Fucking up the language isn’t worth 0.3% on some stupid benchmark, period.
> Modern languages like Java or Swift just give this well defined semantics and call it a day because this isn’t a good hill to die on. Fucking up the language isn’t worth 0.3% on some stupid benchmark, period.
I saw some projects just opt to use -fwrapv as a gcc option. clang I think has the same one too now. The docs for gcc mentions [1] "This flag enables some optimizations and disables others"
[1] https://gcc.gnu.org/onlinedocs/gcc/Code-Gen-Options.html
- signed overflow was defined to wrap like -fwrapv does.
- overflow of any kind is guaranteed to trap.
The latter would be better for sanity and security but I’ll take what I can get.
(what i've usually done for hunting down coroutine errors is use a tracing profiler, or i guess just printf debugging. it is probably the worst part still, tho)
I crave the ergonomy of rust development. I use Rust at my job (not game dev) and it sucks to switch back to C++ for my side projects
But I resist for the moment, because I fear it won't be easy as I predict and it would delay my projects.
I already started using this list of features and refactored most of my code for c++20. I hope C++ will continue on that path and catch up Rust. But there are still so many things missing
In the meantime I refactor little by little my C++ projects to be "rust ready": hierchical ownership, data oriented with minimalist oop. So the day I can't resist no more I will be able to quickly rewrite it in Rust
If you either commit to dylib or C++'s DLLs, you still have to recompile everything on version change, unless you also use the C ABI in C++.
This is only for development. Shipping builds will usually statically link everything.
But, as far as I understand, the boundary layer has to still be C (the side that loads DLLs and stuff), because of the natural limitations of templated languages and linkers.
And as soon as you change any interface you'd need to recompile more parts of the code. The same can be applied with Rust using dylib. At the end, the glue code always end up being C.
It's easier to have a long term cross version cross compiler stable C ABI, but if you're talking a single toolchain that simplifies the problem tremendously and you can absolutely do that with C++ in practice at that point
If you don't care about the ABI being stable then you can use a Rust-based ABI, but you're essentially just static linking everything then. Not sure how that works out for the game developers.