CppCon 2022 Best Practices Every C++ Programmer Needs to Follow – Oz Syed
isocpp.org
isocpp.org
The level is "count your new and delete and see that you have the same amount".
https://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines...
I need to work on a large C++ code base soon and haven’t in 2-3 years, and tbh have neglected leveraging such tool properly in the past, so hoping for some good pointers on a good setup on Linux.
https://www.incredibuild.com/blog/top-9-c-static-code-analys...
I would imagine that nowadays you should use new/delete only if you have to account for every CPU cycle to literally get every drop of performance...
But anyway, if performance matters at all you shouldn't heap-allocate in any hot code path in the first place. E.g. if memory management overhead is showing up in profiling, don't start looking for a faster general-purpose allocator, but instead rethink your memory management strategy ... and "I'll just put every C++ object behind a smart pointer" isn't really a memory management strategy ;)
Isn’t it usually “use raw pointers, if you have to”?
So, "use smart points, if you have to", should have been: "if you need to allocate memory dynamically, use smart pointers".
Oooh. Okay. :p I was like, "As opposed to what, raw pointers?!"
Technically this works and it's reasonably safe, because RAII takes care of memory management, just like GC takes care of memory management in Java.
But the problem there is deeper: it's not smart pointers vs raw pointers, but that each small object lives in its own heap allocation.
That's the typical scenario where you will see memory management becoming a performance problem (for at least two reasons: (1) heap allocation and especially deallocation isn't free, and (2) lots of cache misses when accessing objects that are spread more or less randomly over the address space).
Ideally you'd group many similar objects into long-lived and compact arrays, process the data in those arrays in tight loops, minimize pointer indirections, and minimize allocations (ideally you'd only allocate those arrays once at program start, or when those arrays need to grow). Once you did all those things, memory managament suddenly becomes a non-problem (because you only have a handful of long-lived allocations to care about), and RAII becomes a lot less useful (because the 'objects' in those arrays will most likely just be C-style plain-old-data structs). This is basically the 'antithesis' to OOP. And before you know it you're back to writing C code, like it happened to me ;P
TL;DR: there is no silver bullet for automatic memory management if performance matters.
Are you saying smart pointers and there overhead make the heap allocations even bigger and that’s the cause of the cache misses?
> But the problem there is deeper: it's not smart pointers vs raw pointers, but that each small object lives in its own heap allocation.
Smart-pointers do encourage to create small heap allocations though (because unlike raw pointers they manage ownership).
IMHO from gamedev: HN is not going to get it ever. They either have moved to high level languages or do not need performance at all.
1. https://sean-parent.stlab.cc/presentations/2013-09-11-cpp-se...
I find it annoying when people say you “need” some best practice. The things you actually need are generally enforced by the tools (compiler errors in this case?). Everything else is subjective and/or depends on your use case.
I haven't looked at this specific list yet, but I am yet to see a C++ project in which code works just by following what the tools tell you.
No. This is specifically C++ which has IFNDR ("Ill-formed, No Diagnostic Required") which means the ISO document says there are things (a lot of things it turns out, WG21 appears to have given up even trying to enumerate them) which aren't valid C++ - and so you mustn't do them because the resulting program is meaningless - and yet the compiler isn't expected to detect them and reject your program.
Henry Gordon Rice wrote an important PhD thesis about computation in like 1951. Rice's Theorem says that all non-trivial semantic properties are Undecidable. But we want semantic properties! So, if your programming language is going to have non-trivial semantic properties then you have two practical options:
1. Accept all programs which may have the desired semantic properties, since you can't always decide which those are, you give up and accept programs where you aren't sure, these programs are nonsense, but too bad. That's C++ with IFNDR
2. Reject all programs which may not have the desired semantic properties, since you can't always decide which those are either, you give up and spit out a compiler error when you aren't sure. If a program is rejected maybe the human programmer will rewrite it so that it's acceptable, which is a burden on them. That's what Rust does.
For example maybe when mainprogram.cpp is compiled, the variable max_princesses is defined as the literal 10, but when the princess.cpp file is compiled, perhaps several minutes later, the variable max_princesses is now defined as the literal 0.5 - that's not even the same type! What happens? The ODR means that's IFNDR, so instead of C++ needing to somehow guarantee that this definitely is caught by compilers [these days some compilers will catch some ODR violations but it's not all of them and not always] the standard just says too bad, that's not a valid C++ program so whatever it does is your problem.
A really shiny modern example of IFNDR is C++ 20 Concepts semantic requirements. See, functionally C++ 20 Concepts are just syntax matching, but their names imply semantic value, the standard says they do have semantic value, but it's not actually enforced by the tooling, so syntactically float (a floating point number) matches the concept std::totally_ordered - but of course floats aren't actually totally ordered, that's silly. How do they square this circle? IFNDR. Using a type which matches the syntax, but doesn't fulfil the semantic criteria means your C++ program is ill-formed, but the compiler wasn't expected to tell you about that, your program is just gibberish, it might do anything - after all, floats aren't in fact a totally ordered type.
So of your two examples, here's what I would expect. In the first one, I would expect that either everything would work, or I'd get a syntax error, depending on which definition was actually used at the time I compiled the line in question. Are there examples where anything else happens besides those two options? (OK, I guess there's also the option that two different compilation units have two different definitions at the time I compile them, and then I link them together. Best case the linker catches it; worst case I'm doing floating point operations on what is sometimes an integer variable, and I can see some very weird things happening from there.)
And in the second example, I'd expect everything to work right up until I tried to sort (or whatever) the floating point numbers, at which point it would either work, or fail to sort, or infinite loop, depending on the exact floating point values that the program was operating on. Here I could see there being other options, depending on exactly what the program was trying to do, but not dramatically different. And, would it do anything but work as expected if there were no NANs or negative zeroes or something exotic like that?
That is: Despite the "it can do anything" statements, in practice, with production compilers, does it do completely unreasonable things? Or does the "anything" it does have some reasonableness to it?
Yeah, I know, I'm not guaranteed that. In practice, I don't actually care how my code might break on Windows 3000 with its new 197-bit bytes. I care some about problems on platforms and compilers that it's reasonably likely to need to run on someday. (I'd care more if I were writing library code, and even more if I were writing code for the STL. But I'm not, and while portability is desirable, it's not the only input to decisions.)
The compilers aren't intentionally breaking your code, but they have no responsibility for what happens once you break the rules. I care about Correctness too much for this to be acceptable. If I wrote 100 programs and ten are faulty, I'd rather have twenty errors, half of which are false positives and half catch my ten mistakes, than five errors and five of the "working" programs compile but have mysterious bugs because they're faulty.
It's 5 minutes of very basic and generic statements, like "don't forget to free your memory", "follow rule of five" and "test your code".
Like engineering, methodology an design are always tradeoffs, highly dependent on context.
Business constraints, company culture, legacy code, team strength and even individual preferences, all of that should be taken into account, there are different appropriate styles for different cases.
The only "best practice" is a wide knowledge of what can be done, and the ability to pick and mix to optimize given the context.
Best practices at a minimum can be guidelines for how to build, and at there best they describe exactly what you should do based on the trade-offs and your current situation.
The coding style that is used in safety critical embedded software or high frequency trading would not be appropriate for gamedev or web, and vice versa.
Sometimes the "good reason" is that a certain way won't work with the giant code base that already exists and the current way isn't actively causing any bug or major security risks. It can be hard to convince team members that want everything to be the "best" way of that, but that's a problem with people, not a problem with the concept of "Best Practices".
Every major life-or-death software system has "Best Practices". Hopefully, everyone that works on those system understands it's a fluid concept that will change as new things are learned but it's important to understand and follow them unless there's a valid, peer reviewed, reason not to.
So, I would say yes, it is worth. But truly, it probably depends on what you're coding.
One strategy is to put a minimal application-specific C layer around your C++ library, and call that from the Rust app. I’ve done that in the past and it kept the Rusties on the team happy enough. But you need to be fairly certain that the API can remain small enough and high-level enough even as needs grow, because nobody wants to be maintaining a C wrapper of hundreds of API calls a few years down the road.
In C++ circles, safety conscious developers are always fighting against C like code, or disabling bounds checking.
If you are in a team that values bounds checking enabled by default, and where C style coding is forbidden unless for FFI reasons, then it is another matter.
See Jason Turner's starter packs.
Basically it isn't perfect as solution, but by enabling them as if they part of the language, we get to clean up wrong defaults, and force best practices, specially if they are integrated as part of a CI/CD pipeline, including breaking the build when they aren't followed upon.
It isn't a fullproof solution, but better than nothing, and plenty of domains aren't going to switch to anything else anyway.
It is kind of ironic that even C's authors saw the need of such tooling and created lint in 1979, but apparently the FOSS hardly cared about such kind of tooling until clang came to be.
This is an example of an eggcorn, humans acquire language by exposure, you misunderstand a word or saying, you analyse its apparent meaning based on what you thought you understood, and then you apply that analysis, it's a completely normal part of human language use. In some cases, eggcorns become normalised enough that it's reasonable to say they're just a variant use, but in most cases they'd be regarded as errors, and so it's probably useful to know if you've picked up any eggcorns. https://en.wikipedia.org/wiki/Eggcorn
It's true that Rust does not have a written specification that clearly delineates what is and isn't UB in a single place. But:
1. UB is impossible in safe code (modulo bugs in unsafe code)
2. There are resources such as the Rustinomicon (https://doc.rust-lang.org/nomicon/) that provide a detailed guide on what is and isn't allowed in unsafe code.
In practice, it's much easier to avoid UB in Rust than it is in C++.
Based on that definition it feels like it should be possible to have UB outside of memory violations, is there really no UB in languages like Java/Haskell/Go?
And I am assuming something like the NullPointerException comes with a huge performance hit? Otherwise I assume every systems language would do something similar.
Presumably it's not defined because the behaviour depends on the signedness representation.
Since we can be sure if it ever happens your code has a bug, making it undefined is a good thing: the compiler can then assume it doesn't happen and so back track to prove some other things can't happen and so make your program run a little faster.
But C++ also isn't perfect, there are plenty of programs for which no two compiler developers can agree on whether they have UB. The C++ spec language is just too ambiguous and underspecified in several areas.
If you want to be sure, you need an actual machine-checkable formal specification. Neither C++ nor Rust have that.
In the end, what really matter is the contract between the programmer and the compiler: are compilers allowed to break a program in weird ways because the programmer forgot about one of the arcane rules in the spec? For C++ and unsafe Rust, the answer is yes (we don't know how to build optimizing compilers for low-level languages otherwise). But for safe Rust, the answer is no. That's a big deal.
The reality is that C++ can be used very well and very badly. What makes C++ code good or not is not defined by the fact it's more or less C++, but rather by who wrote it.
imho Rust is modern C++ done right (i.e. without all the baggage from the C++ past due to backwards compatibility) -> less mental overhead, still safer, more convenient features, less debugging, proper tooling...
To do C++ right, you'd need to do "marriage of OOP and low level programming" right, which is probably impossible because it's the most cursed combination of ideas ever conceived.
WAIT, I take it back -- D improves on C++, so it's "C++ done right". :D
not really, C++ is a multiparadigm language, so is Rust
Rust makes so many guarantees that I expect some of that Ph.D research will start to turn into concrete language improvements in ways it never could before. It'll be very exciting to watch the next 5-10 years of Rust as the best overlooked research ideas start to have a viable path to production projects.