The Performance Impact of C++'s `final` Keyword
16bpp.net
16bpp.net
Inlining has other requirements as well -- LTO pretty much covers it.
The article doesn't have sufficient data to tell whether the testcase is built in such a way that any of these optimizations can happen or is beneficial.
$ cat animal.h cat.cpp main.cpp
// animal.h
#pragma once
class animal {
public:
virtual ~animal() {}
virtual void speak() = 0;
};
animal& get_mystery_animal();
// cat.cpp
#include "animal.h"
#include <cstdio>
class cat final : public animal {
public:
~cat() override{}
void speak() override{
puts("meow");
}
};
static cat garfield{};
animal& get_mystery_animal() {
return garfield;
}
// main.cpp
#include "animal.h"
int main() {
animal& a = get_mystery_animal();
a.speak();
}
$ make clean && CXX=clang++ make -j && objdump --disassemble=main -C lto_test
rm -f *.o lto_test
clang++ -c -flto -O3 -g cat.cpp -o cat.o
clang++ -c -flto -O3 -g main.cpp -o main.o
clang++ -flto -O3 -g cat.o main.o -o lto_test
lto_test: file format elf64-x86-64
Disassembly of section .init:
Disassembly of section .plt:
Disassembly of section .plt.got:
Disassembly of section .text:
00000000000011b0 <main>:
11b0: 50 push %rax
11b1: 48 8b 05 58 2e 00 00 mov 0x2e58(%rip),%rax # 4010 <garfield>
11b8: 48 8d 3d 51 2e 00 00 lea 0x2e51(%rip),%rdi # 4010 <garfield>
11bf: ff 50 10 call *0x10(%rax)
11c2: 31 c0 xor %eax,%eax
11c4: 59 pop %rcx
11c5: c3 ret
Disassembly of section .fini:
$ make clean && CXX=g++ make -j && objdump --disassemble=main -C lto_test|sed -e 's,^, ,'
rm -f *.o lto_test
g++ -c -flto -O3 -g cat.cpp -o cat.o
g++ -c -flto -O3 -g main.cpp -o main.o
g++ -flto -O3 -g cat.o main.o -o lto_test
lto_test: file format elf64-x86-64
Disassembly of section .init:
Disassembly of section .plt:
Disassembly of section .plt.got:
Disassembly of section .text:
0000000000001090 <main>:
1090: 48 83 ec 08 sub $0x8,%rsp
1094: 48 8d 3d 75 2f 00 00 lea 0x2f75(%rip),%rdi # 4010 <garfield>
109b: e8 50 01 00 00 call 11f0 <cat::speak()>
10a0: 31 c0 xor %eax,%eax
10a2: 48 83 c4 08 add $0x8,%rsp
10a6: c3 ret
Disassembly of section .fini:You can tell it "I won't do that" though with additional flags, like Clang's -fwhole-program-vtables, and even then it's not that simple. There was an effort in Clang to better support whole program devirtualization, but I haven't been following what kind of progress has been made: https://groups.google.com/g/llvm-dev/c/6LfIiAo9g68?pli=1
Maybe I can set this option at work. Though it's scary because I'd have to be certain.
Lots of code gets slower if it might need to be called from something not currently in the compiler's scope. That's essentially what ABI overhead is. If there isn't already, there should be a compiler flag that says "this is the whole program, have at it" which implies the vtables option.
* It is possible with `dlopen()` to load code objects that violate the assumptions made during compilation.
* The presence of runtime configuration mechanisms and application input can make it impossible to anticipate things like the choice of implementations of an interface.
One can always strive to reduce such situations, but it might simply not be necessary if a JIT is present.
Is there a theory as to how devirtualisation could hurt performance?
The main advantages to inlining are (1) avoiding a jump and other function call overhead, (2) the ability to push down optimizations.
If you execute the "same" code (same instructions, different location) in many places that can cause cache evictions and other slowdowns. It's worse if some minor optimizations were applied by the inlining, so you have more types of instructions to unpack.
The question, roughly, is whether the gains exceed the costs. This can be a bit hard to determine because it can depend on the size of the whole program and other non-local parameters, leading to performance cliffs at various stages of complexity. Microbenchmarks will tend to suggest inlining is better in more cases that it actually is.
Over time you get a feel for which functions should be inlined. E.g., very often you'll have guard clauses or whatnot around a trivial amount of work when the caller is expected to be able to prove the guarded information at compile-time. A function call takes space in the generated assembly too, and if you're only guarding a few instructions it's usually worth forcing an inline (even in places where the compiler's heuristics would choose not to because the guard clauses take up too much space), regardless of the potential cache costs.
If you have something like a `while` loop and that while loop's instructions fit neatly on the cache line, then executing that loop can be quiet fast even if you have to jump to different code locations to do the internals. However, if you pump in more instructions in that loop you can exceed the length of the cache line which causes you to need more memory loads to do the same work.
It can also create more code. A method that took a `foo(NotFinal& bar)` could be duplicated by the compiler for the specialized cases which would be bad if there's a lot of implementations of `NotFinal` that end up being marshalled into foo. You could end up loading multiple implementations of the same function which may be slower than just keeping the virtual dispatch tables warm.
And if the devirtualisation leads to inlining, that results in code bloat which can lower performance though more instruction cache misses, which are not cheap.
Inlining is actually pretty evil. It almost always speeds things up for microbenchmarks, as such benchmarks easily fit in icache. So programmers and modern compilers often go out of their way to do more inlining. But when you apply too much inlining to a whole program, things start to slow down.
But it's not like inlining is universally bad in larger program, inlining can enable further optimisations, mostly because it allows constant propagation to travel across function boundaries.
Basically, compilers need better heuristics about when they should be inlining. If it's just saving the overhead of a lightweight call, then they shouldn't be inlining.
No it's not. Except if you __force_inline__ everything, of course.
Inlining reduces the number of instructions in a lot of cases. Especially when things are abstracted and factored with lot of indirections into small functions that calls other small functions and so on. Consider a 'isEmpty' function, which dissolves to 1 cpu instruction once inlined, compared with a call/save reg/compare/return. Highly dynamic code (with most functions being virtual) tend to result in a fest of chained calls, jumping into functions doing very little work. Yes the stack is usually hot and fast, but spending 80% of the instructions doing stack management is still a big waste.
Compilers already have good heuristics about when they should be inlining, chances are they are a lot better at it than you. They don't always inline, and that's not possible anyway.
My experience is that compiler do marvels with inlining decisions when there are lots of small functions they _can_ inline if they want to. It gives the compiler a lot of freedom. Lambdas are great for that as well.
Make sure you make the most possible compile-time information available to the compiler, factor your code, don't have huge functions, and let the compiler do its magic. As a plus, you can have high level abstractions, deep hierarchies, and still get excellent performances.
As you say: “chances are they are a lot better at it than you”. Infrequently they are not.
In a moderately-sized codebase I regularly work on, I use __attribute__((noinline)) nearly ten times as often as __attribute__((always_inline)). And I use __attribute__((cold)) even more than noinline.
So yeah, I can kind of see why someone would say inlining is 'evil', though I think it's more accurate to say that it's just not possible for compilers to figure out these kinds of details without copious hints (like PGO).
When writing ultra-robust code that has to survive every vaguely plausible contingency in a graceful way, the code is littered with code paths that only exist for astronomically improbable situations. The branch predictor can figure this out but the compiler frequently cannot without explicit instructions to not pollute the i-cache.
The first is that when building a code base you don't necessarily know what it's being compiled with. And so even if there were a super-amazing compiler, there's no guarantee that's what will be compiling your code. Making it explicit, so long as you have a reasonably good idea of what you're doing, is generally just a good idea. It also conveys intent to some degree, especially things like final.
The second is that I think the saying 'premature optimization is the root of all evil' is the root of all evil. Because that mindset has gradually transitioned to being against optimization in general outside of the most primitive things like not running critical sections in O(N^2) when they could be O(N). And I think it's this mindset that has gradually brought us to where we are today where need what what would have been a literal supercomputer not that long ago, to run a word processor. It's like death by a thousand cuts, and quite ridiculous.
The greater evil is putting a one-sentence quote out of context:
""" There is no doubt that the grail of efficiency leads to abuse. Programmers waste enormous amounts of time thinking about, or worrying about, the speed of noncritical parts of their programs, and these attempts at efficiency actually have a strong negative impact when debugging and maintenance are considered. We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil.
Yet we should not pass up our opportunities in that critical 3%. A good programmer will not be lulled into complacency by such reasoning, he will be wise to look carefully at the critical code; but only after that code has been identified. It is often a mistake to make a priori judgments about what parts of a program are really critical, since the universal experience of programmers who have been using measurement tools has been that their intuitive guesses fail. After working with such tools for seven years, I've become convinced that all compilers written from now on should be designed to provide all programmers with feedback indicating what parts of their programs are costing the most; indeed, this feedback should be supplied automatically unless it has been specifically turned off. """
So referencing something in particular from Unreal Engine, they actually created a caching system for converting between a quaternion and a rotator (euler rotation)! Obviously that sort of conversion isn't going to, in a million years, be even close to a bottleneck. That conversion is quite cheap on modern hardware, and so that caching system probably only gives the engine one of those 0.1% boosts in performance. But there are literally thousands of these "small efficiencies" spread all throughout the code. And it yields a final product that runs dramatically better than comparable engines.
The branch predictors actually hash the history of the last few branches taken into the branch prediction query. So the exact same branch within a child function will map different branch predictors entries depending on which parent function it was called from, and there is no benifit to inlining.
It also means that branch predictor can also learn correlations between branches within a function. Like when a branches at the top and bottom of functions share conditions, or have inverted conditions.
Guarded devirtualization is also cheaper than virtual calls, even when it has to do
if (instance is SpecificType st) { st.Call() }
else { instance.Call() }
or even chain multiple checks at once (with either regular ifs or emitting a jump table)This technique is heavily used in various forms by .NET, JVM and JavaScript JIT implementations (other platforms also do that, but these are the major ones)
The first two devirtualize virtual and interface calls (important in Java because all calls default to virtual, important in C# because people like to abuse interfaces and occasionally inheritance, C# delegates are also devirtualized/inlined now). The JS JIT (like V8) performs "inline caching" which is similar where for known object shapes property access is shape type identifier comparison and direct property read instead of keyed lookup which is way more expensive.
Also you are correct - virtual calls are not terribly expensive, but they encroach on ever limited* CPU resources like indirect jump and load predictors and, as noted in parent comments, block inlining, which is highly undesirable.
[0] https://github.com/dotnet/runtime/blob/5111fdc0dc464f01647d6...
[1] https://github.com/dotnet/runtime/blob/main/docs/design/core... (mind you, the text was initially written 18 years ago, wow)
* through great effort of our industry to take back whatever performance wins each generation brings with even more abstractions that fail to improve our productivity
I think that enabling inlining is just one of the indirect consequences of devirtualization, and perhaps one that is largely irrelevant for performance improvements.
The whole point of devirtualization is eliminating the need to resort to pointer dereferencing when calling virtual members. The main trait of a virtual class is it's use of a vtable that requires dereferencing virtual members to access each and every one of them.
In classes with larger inheritance chains, you can easily have more than one pointer dereferencing taking place before you call a virtual members function.
Once a class is final, none of that is required anymore. When a member is referred, no dereferencing takes place.
Devirtualization helps performance because you are able to benefit from inheritance and not have to pay a performance penalty for that. Without the final keyword, a performance oriented project would need to be architected to not use inheritance at all, or in the very least in code in the hot path, because that sneaks gratuitous pointer dereferences all over the place, which require running extra operations and has a negative impact on caching.
The whole purpose of the final keyword is that compilers can easily eliminate all pointer dereferencing used by virtual members. What stops them from applying this optimization is that they have no information on whether that class will be inherited and one of its members will either override any of its members or invoke any member function implemented by one of its parent classes.
With the introduction of the final keyword, you are now able to tell the compiler "from thereon, this is exactly what you get" and the compiler can trim out anything loose.
Inlining is by far the most impactful optimization here, because it can eliminate the call altogether, and thus specialize the called function to the callsite, lifting constants, hoisting loop variables, etc.
My guess is this is why he didn't see any speedup: all the code could fit inside the L2 cache, so he did not have to pay for RAM access for the deference.
The number of different classes is important, not the number of objects as they have the same small number of vtable pointers.
It might be different for large codebases like Chrome and Firefox.
I was going to eliminate polymorphism altogether for this object but later figured out how to refactor so that this particular call could be called once a millisecond. Then if more work was needed, it would dispatch a task to a dedicated CPU.
This was an incredibly performant improvement which made a significant difference to my P&L.
In general if you're manipulating values that fit into registers and work on a platform with a shitty ABI,you need to be very careful of what your function call boundaries look like.
The most obvious example is SIMD programming on Windows x86 32-bit.
If they cannot be predicted, write your code accordingly.
Of course you have to worry about pointer chasing, when you can easily avoid it. Either via a switch to a single indirection (by passing method pointers around) or inlining with final. Or other compile-time specialization.
In general it takes a significant amount of nondeterministic pointer chasing to fool modern branch predictors. Decades of research have been put into optimizing the hardware for languages like C++ and Java, both of which exhibit a lot of pointer chasing.
Though, that assumes a correct prediction. But modern branch predictors are really good, they can track and correctly predict hundreds (if not thousands) of indirect calls, taking into account the history of the last few branches (so it can even get an idea of what class is currently being executed, and make branch predictions based on that). Modern branch predictors do a really good job at chewing up indirect branches in hot sequences of code.
Virtual functions are probably the most harmful for warm code. We are talking about code that's executed too often to be considered cold code, but not often enough to stick around in the branch predictors' cache, executed only a few hundred times a second. It's a death by a thousand cuts type thing. And that's where devirtualisation will help the most...
As long as you don't go too far with the inlining and start causeing icache misses with code bloat. In an ideal would the compiler would inline enough to devirtualise the class, but not necessarily inline the actual function (unless they are small, or only called from one place)
virtual inheritance. Regular old inheritance does not need or benefit from devirtualization. This is why the CRTP exists.
CRTP does not exist for that. CRTP was one of the many happy accidents in template metaprogramming that happened to be discovered when doing recursive templates.
Also, you've missed the whole point. CRTP is a way to rearchitect your code to avoid dereferencing pointers to virtual members in inheritance. The whole point is that with final you do not need to pull tricks: just tell the compiler that you don't want the class to be inherited, and the compiler picks up from there and does everything for you.
Please read my post. That's not my claim. I think I was very clear.
What you're talking about is dynamic dispatch
This is not a thing in C++; vtables are flat, not nested. Function pointers are always 1 dereference away.
Funny how things work. From working with Julia I've built a good intuition for guessing when functions would be inlined. And yet, I've never heard the word devirtualization until now.
Quick example, I got in an argument with someone a few years ago that claimed in C# that a `switch` was better than an `if(x==1) elseif(x==2)...` because switch was "faster" and rejected my PR. I mentioned that that doesn't appear to be true, we went back and forth until I did a compile-then-decompile of a minimal test with equality-based-ifs, and showed that the compiler actually converts equality-based-ifs to `switch` behind the scenes. The guy accepted my PR after that.
But there's tons of this stuff like this in CS, and I kind of blame professors for a lot of it [1]. A large part of becoming a decent engineer [2] for me was learning to stop trusting what professors taught me in college. Most of what they said was fine, but you can't assume that; what they tell you could be out of date, or simply never correct to begin with, and as far as I can tell you have to always test these things.
It doesn't help that a lot of these "it's faster" arguments are often reductive because they only are faster in extremely minimal tests. Sometimes a microbenchmark will show that something is faster, and there's value in that, but I think it's important that that can also be a small percentage of the total program; compilers are obscenely good at optimizing nowadays, it can be difficult to determine when something will be optimized, and your assertion that something is "faster" might not actually be true in a non-trivial program.
This is why I don't really like doing any kind of major optimizations before the program actually works. I try to keep the program in a reasonable Big-O and I try and minimize network calls cuz of latency, but I don't bother with any kind of micro-optimizations in the first draft. I don't mess with bitwise, I don't concern myself on which version of a particular data structure is a millisecond faster, I don't focus too much on whether I can get away with a smaller sized float, etc. Once I know that the program is correct, then I benchmark to see if any kind of micro-optimizations will actually matter, and often they really don't.
[1] That includes me up to about a year ago.
[2] At least I like to pretend I am.
Writing well structured readable code is typically far more important than making it twice as fast. And those times can rarely be predicted beforehand, so you should mostly not worry about it until you see real performance problems.
For example, much to the annoyance of a lot of people, I don't typically use floating point numbers when I start out. I will use the "decimal" or "money" types of the language, or GMP if I'm using C. When I do that, I can be sure that I won't have to worry about any kind of funky overflow issues or bizarre rounding problems. There might be a performance overhead associated with it, but then I have to ask myself "how often is this actually called?"
If the answer is "a billion times" or "once in every iteration of the event loop" or something, then I will probably eventually go back and figure out if I can use a float or convert it to an integer-based thing, but in a lot of cases the answer is "like ten or twenty times", and at that point I'm not even 100% sure it would be even measurable to change to the "faster" implementations.
What annoys me is that people will act like they really care about speed, do all these annoying micro-optimizations, and then forget that pretty much all of them get wiped out immediately upon hitting the network, since the latency associated with that is obscene.
Why should you care about performance?
I can give you my personal experience: I’ve been working on a Java web/application server for the past 15 years and a typical request (only reading, not writing to the db) would take maybe 4-5 ms to execute. That includes HTTP request parsing, JSON parsing, session validation, method execution, JSON serialization, and HTTP response dispatch. Over the past 9 months I have refactored the entire application for performance and a typical request now takes about 0.25 ms or 250 microseconds. The computer is doing so much less work to accomplish the same tasks, it’s almost silly how much work it was doing before. And the result is the machine can handle 20x more requests in the same amount of time. If it could handle 200 requests per second per core before, now it can handle 4000. That means the need to scale is felt 20x less intensely, which means less complexity around scaling.
High performance means reduced scaling requirements.
I'm not saying you completely throw caution to the wind, I'm just saying that there's a finite amount of human resources and it can really vary how you want to allocate them. Sometimes the better path is to just throw money at the problem.
It really depends.
I personally don’t like the idea of throwing compute at a slow solution. I like when the extra effort has been put into something. The good feeling I get from interacting with something that is optimal or excellent is an end in itself and one of the things I live for.
CPUs are ridiculously fast now, and compilers are really really good now too. I'm not going to say that processing speed is a "solved" problem, but I am going to say that in a lot of performance-related cases the CPU processing is probably not your problem. I will admit that this kind of pokes holes in my previous response, because introducing more machines into the mix will almost certainly increase latency, but I think it more or less holds depending on context.
But I think it really is a matter of nuance, which you hinted at. If I'm making an admin screen that's going to have like a dozen users max, then a slow, crappy solution is probably fine; the requests will be served fast enough to where no one will notice anyway, and you can probably even get away with the cheapest machine/VM. If I'm making an FPS game that has 100,000 concurrent users, then it almost certainly will be beneficial to squeeze out as much performance out of the machine as possible, both CPU and latency-wise.
But as I keep repeating everywhere, you have to measure. You cannot assume that your intuition is going to be right, particularly at-scale.
(Ironically, HN is buckling under load right now, or some other issue.)
If your problem can fit on one server, it can massively reduce engineering and infrastructure costs.
It is why many language ecosystems suffered from performance issues for a really long time even if completely unwarranted.
Is changing ifs to switch or vice versa, as outlined in the post above, a waste of time? Yes, unless you are writing some encoding algorithm or a parser, it will not matter. The compiler will lower trivial statements to the same codegen and it will not impact the resulting performance anyway even if there was difference given a problem the code was solving.
However, there are things that do cost like interface spam, abusing lambdas writing needlessly complex wokflow-style patterns (which are also less readable and worse in 8 out of 10 instances), not caching objects that always have the same value, etc.
These kinds of issues, for example, plagued .NET ecosystem until more recent culture shift where it started to be cool once again to focus on performance. It wasn't being helped by the notion of "well-structured code" being just idiotic "clean architecture" and "GoF patterns" style dogma applied to smallest applications and simplest of business domains.
(it is also the reason why picking slow languages in general is a really bad idea - everything costs more and you have way less leeway for no productivity win - Ruby and Python, and JS with Node.js are less productive to write in than C#/F#, Kotlin/Java or Go(under some conditions))
There are plenty of cases where even the "slow" implementation is more than fast enough, and there are also plenty of cases where the "correct" solution (from a big-O or intuition perspective) is actually slower than the dumb case. Intuition helps, you have to measure and/or look at the compiled results if you want to ensure correct numbers.
An example that really annoys me is how every whiteboard interview ends up being "interesting ways to use a hashmap", which isn't inherently an issue, but they will usually be so small-scoped that an iterative "array of pairs" might actually be cheaper than paying the up-front cost of hashing and potentially dealing with collisions. Interviews almost always ignore constant factors, and that's fair enough, but in reality constant factors can matter, and we're training future employees to ignore that.
I'll say it again: as far as I can tell, you have to measure if you want to know if your result is "faster". "Measuring" might involve memory profilers, or dumb timers, or a mixture of both. Gut instincts are often wrong.
that said, c++ is usually a language you use when you care about performance, at least to an extent. it's worth understanding features like nrvo and rewriting functions to allow the compiler to pick the optimization if it doesn't hurt readability too much.
That said, the linear test is often faster due to CPU caches, which is why JITs will often convert switches to if/elses.
IMO, switch is clearer in general and potentially faster (at very least the same speed) so it should be preferred when dealing with 3+ if/elseif statements.
if statements are dumber, and maybe arguably uglier, but I feel like they're also more clear, and people don't try and be clever with them.
For example, with java there's enhanced switch that looks like this
var val = switch(foo) {
case 1, 2, 3 -> bar;
case 4 -> baz;
default -> {
yield bat();
}
}
The C style switch break stuff is definitely a language mistake.However most languages have pretty permissive switch statements just like C.
So no offense, but I would revisit the wider world of language constructs before claiming that switch statements are "all bad". There are plenty of bad languages or languages with poor implementations of syntax, that do not make the fundamental language construct bad.
var len = slice switch
{
null => 0,
"Hello" or "World" => 1,
['@', ..var tags] => tags.Length,
['{', ..var body, '}'] => body.Length,
_ => slice.Length,
};
(it supports a lot more patterns but that wouldn't fit)However, the moment you add a side effect or something more complicated like a method call, it becomes really hard for the complier to know if that sort of optimization is safe to do.
The benefit of the switch statement is that it's already well positioned for the compiler to optimize as it does not have the "you must run these evaluations in order" requirement. It forces you to write code that is fairly compiler friendly.
All that said, probably a waste of time debating :D. Ideally you have profiled your code and the profiler has told you "this is the slow block" before you get to the point of worrying about how to make it faster.
Awkward switch syntax aside, the switch is simpler to reason about. Fundamentally we should strive to keep our code simple to understand and verify, not worry about compiler optimizations (on the first pass).
It turns out the algorithmic complexity of a switch statement and the equivalent series of if-statements is identical. The bijective mapping between them is close to the identity function. Does a naive compiler exist that doesn't emit the same instructions for both, at least outside of toy hobby project compilers written by amateurs with no experience?
If statements are unbounded, unconstrained logic constructs, whereas switch statements are type-checkable. The concern about missing break statements here is irrelevant, where your linter/compiler can warn about missing switch cases they can easily warn about non-terminated (non-explicitly marked as fall-through) cases.
For non-compiled languages (so branch prediction is not possible because the code is not even loaded), switch statements also provide a speed-up, i.e. the parser can immediately evaluate the branch to execute vs being forced to evaluate intermediate steps (and the conditions to each if statement can produce side-effects e.g. if(checkAndDo()) { ... } else if (checkAndDoB()) { ... } else if (checkAndDoC()) { ... }
Which, of course, is a potential use of if statements that switches cannot use (although side-effects are usually bad, if you listened to your CS profs)... And again a sort of "static analysis" guarantee that switches can provide that if statements cannot.
But I agree, algorithmic complexity is generally the only thing I focus on, and even then it's almost always a case of "will that actually matter?" If I know that `n` is never going to be more than like `10`, I might not bother trying to optimize an O(n^2) operation.
What I feel often gets ignored in these conversations is latency; people obsess over some "optimization" they learned in college a decade ago, and ignore the 200 HTTP or Redis calls being made ten lines below, despite the fact that the latter will have a substantially higher impact on performance.
My experience is the opposite - a sizeable chain of ifs has more that can go wrong precisely because it is more flexible. If I'm looking at a switch, I immediately know, for instance, that none of the tests modifies anything.
Meanwhile, while a missing break can be a brutal error in a language that allows it, it's usually trivial to set up linting to require either an explicit break or a comment indicating fallthrough.
This is not entirely true either... Measure. There are many cases where the optimiser will vectorise a certian algorithm but not another... In many cases On^2 vectorised may be significantly faster than On or Onlogn even for very large datasets depending on your data...
Make your algorithms generic and it won't matter which one you use, if you find that one is slower swap it for the quicker one. Depending on CPU arch and compiler optimisations the fastest algorithm may actually change multiple times in a codebases lifetime even if the usage pattern doesn't change at all.
Reminds me of the classic https://stackoverflow.com/questions/24848359/which-is-faster...
When talking about not assuming optimizations...
32bit float is slower than 64bit float on reasonable modern x86-64.
The reason is that 32bit float is emulated by using 64bit.
Of course if you have several floats you need to optimize against cache.
SIMD/MIMD will benefit of working on smaller width. This is not only true because they do more work per clock but because memory is slow. Super slow compared to the cpu. Optimization is alot about cache misses optimization.
(But remember that the cache line is 64 bytes, so reading a single value smaller than that will take the same time. So it does not matter in theory when comparing one f32 against one f64)
x86-64 requires the hardware to support SSE2, which has native single-precision and double-precision instructions for floating-point (e.g., scalar multiply is MULSS and MULSD, respectively). Both the single precision and the double precision instructions will take the same time, except for DIVSS/DIVSD, where the 32-bit float version is slightly faster (about 2 cycles latency faster, and reciprocal throughput of 3 versus 5 per Agner's tables).
You might be thinking of x87 floating-point units, where all arithmetic is done internally using 80-bit floating-point types. But all x86 chips in like the last 20 years have had SSE units--which are faster anyways. Even in the days when it was the major floating-point units, it wasn't any slower, since all floating-point operations took the same time independent of format. It might be slower if you insisted that code compilation strictly follow IEEE 754 rules, but the solution everybody did was to not do that and that's why things like Java's strictfp or C's FLT_EVAL_METHOD were born. Even in that case, however, 32-bit floats would likely be faster than 64-bit for the simple fact that 32-bit floats can safely be emulated in 80-bit without fear of double rounding but 64-bit floats cannot.
[0] https://gist.github.com/dosshell/495680f0f768ae84a106eb054f2...
Sorry for the confusion and spreading false information.
I've solved a lot of arguments with godbolt and simple performance tests. Some topics are recurring themes among software engineers e.g.:
- compilers are almost always better at micro-optimizations than you are
- disk I/O is almost never a bottleneck in competent designs
- brute-force sequential scans are often optimal algorithms
- memory is best treated as a block device
- vectorization can offer large performance gains
- etc...
No one is immune to this. I am sometimes surprised at the extent to which assumptions are no longer true when I revisit optimization work I did 10+ years ago.
Most performance these days is architectural, so getting the initial design right often has a bigger impact than micro-optimizations and localized Big-O tweaks. You can always go back and tweak algorithms or codegen later but architecture is permanent.
(the techniques that used to work were similar to earlier Java versions and overall very dynamic languages with some exceptions, the techniques that still work and now are required today are the same as in C++ or Rust)
There are a few articles on msft devblogs that cover from-netframework migration to older versions (Core 3.1, 5/6/7):
- https://devblogs.microsoft.com/dotnet/bing-ads-campaign-plat...
- https://devblogs.microsoft.com/dotnet/microsoft-graph-dotnet...
- https://devblogs.microsoft.com/dotnet/the-azure-cosmos-db-jo...
- https://devblogs.microsoft.com/dotnet/one-service-journey-to...
- https://devblogs.microsoft.com/dotnet/microsoft-commerce-dot...
The tl;dr is depending on codebase the latency reduction was anywhere from 2x to 6x, varying per percentile, or the RPS was maintained with CPU usage dropping by ~2-6x.
Now, these are codebases of likely above average quality.
If you consider that moving 6 -> 8 yields another up to 15-30% on average through improved and enabled by default DynamicPGO, and if you also consider that the average codebase is of worse quality than whatever msft has, meaning that DPGO-reliant optimizations scale way better, it is not difficult to see the 10x number.
Keep in mind that while particular regular piece of enterprise code could have improved within bounds of "poor netfx codegen" -> "not far from LLVM with FLTO and PGO", the bottlenecks have changed significantly where previously they could have been in lock contention (within GC or user code), object allocation, object memory copying, e.g. for financial domains - anything including possibly complex Regex queries on imported payment reports (these alone have now difference anywhere between 2 and >1000[0]), and for pretty much every code base also in interface/virtual dispatch for layers upon layers of "clean architecture" solutions.
The vast majority of performance improvements (both compiler+gc and CoreLib+frameworks), which is difficult to think about, given it was 8 years, address the above first and foremost. At my previous employer the migration from NETFX 4.6 to .NET Core 3.1, while also deploying to much more constrained container images compared to beefy Windows Server hosts, reduced latency of most requests by the same factor of >5x (certain request type went from 2s to 350ms). It was my first wow moment when I decided to stay with .NET rather than move over to Go back then (was never a fan of syntax though, and other issues, which subsequently got fixed in .NET, that Go still has, are not tolerable for me).
[0] Cumulative of
https://devblogs.microsoft.com/dotnet/regex-performance-impr...
https://devblogs.microsoft.com/dotnet/regular-expression-imp...
https://devblogs.microsoft.com/dotnet/performance-improvemen...
All of the 6x performance improvement cases seem to be related to using the .net based Kestrel web server instead of IIS web server, which requires marshalling and interprocess communication. Several of the 2x gains appear to be related to using a different database backend. Claims that regex performance has improved a thousand-fold.... seem more troubling than cause for celebration. Were you not precompiling your regex's in the older code? That would be a bug.
Somewhere in there, there might be 30% improvements in .net codegen (it's hard to tell). Profile Guided Optimization (PGO) seems to provide a 35% performance improvement over older versions of .net with PGO disabled. But that's dishonest. PGO was around long before .net Core. And claiming that PGO will provide 10x performance because our code is worse than Microsoft's code insults both our code and our intelligence.
With a Roslyn-based compiler at work I saw 20 % perf improvement just by switching from .NET Core 3.1 to .NET 6. No idea how slow .NET Framework was, though. I probably can't target the code to that anymore.
But for regex even with precompilation, the compiler got a lot better at transforming the regex into an equivalent regex that performs better (automatic atomic grouping to reduce unnecessary backtracking when it's statically known that backtracking won't create more matches for example) and it also benefits a lot from the various vectorized implementations of Index of, etc. Typically with each improvement of one of those core methods for searching stuff in memory there's a corresponding change that uses them in regex.
So where in .NET Framework a regex might walk through a whole string character by character multiple times with backtracking it might be replaced with effectively an EndsWith and LastIndexOfAny call in newer versions.
(the distinction becomes important for targets serviced by Mono, so to outline the difference Mono is usually specified, while CoreCLR and RyuJIT may not be, it also doesn't help that JIT, that is, the IL to machine code compiler, also services NativeAOT, so it gets more annoying to be accurate in a conversation without saying the generic ".net compiler", some people refer to it as JIT/ILC)
And indeed, on the C# -> IL side there's little that's being actually optimized. Besides collection literals there's also switch statements/expressions over strings, along with certain pattern matching constructs that get improved on that side.
Is it a public project?
Also, IIS hosting through Http.sys is still an option that sees separate set of improvements, but that's not relevant in most situations given the move to .NET 8 from Framework usually also involves replacing Windows Server host with a Linux container (though it works perfectly fine on Windows as well).
On Regex, compiled and now source generated automata has seen a lot of work in all recent releases, it is night and day to what it was before - just read the articles. Previously linear scans against heavy internal data structures (matching by hashset) and heavy transient allocations got replaced with bloom-filter style SIMD search and other state of the art text search algorithms[0], on a completely opposite end of a performance spectrum.
So when you have compiler improvements multiplied by changes to CoreLib internals multiplied by changes to frameworks built on top - it's achievable with relative ease. .NET Framework, while performing adequately, was still that slow compared to what we got today.
[0] https://github.com/dotnet/runtime/tree/main/src/libraries/Sy...
And you have misrepresented the contents of the blogs. The projects discussed in the blogs are typically claiming ~30% improvements (perhaps because they weren't using static PGO in their 4.7.0 incarnation), with two dramatic outliers that seem to be related to migrating from IIS to Kestrel.
It’s also convenient to ignore the rest of the content at the links but it seems you’re more interested in proving your argument so the data I provided doesn’t matter.
> Were you not precompiling your regex's in the older code? That would be a bug.
I never heard of this before. Perl has legendary fast regexen and I never heard of this feature. Does Java do it? I don't think so, and the regexes are fast enough in my experience. Can you name a language when regexen are precompiled?When I'm building stuff I try my best to focus on "correctness", and try to come up with an algorithm/design that will encompass all realistic use cases. If I focus on that, it's relatively easy to go back and convert my `decimal` type to a float64, or even convert an if statement into a switch if it's actually faster.
When I was taught about performance, it was all about benchmarking and profiling. I never needed to trust what my professors taught, because they taught me to dig in and find the truth for myself. This was taught alongside the big-O stuff, with several examples where "fast" algorithms are slower on small inputs.
JVM ecosystem has IntelliJ Idea profiler and similar advanced tools (AFAIK).
.NET has VS/Rider/dotnet-trace profilers (they are very detailed) to produce flamegraphs.
Then there are native profilers which can work with any AOT compiled language that produces canonically symbolicated binaries: Rust, C#/F#(AOT mode), Go, Swift, C++, etc.
For example, you can do `samply record ./some_binary`[0] and then explore multi-threaded flamegraph once completed (I use it to profile C#, it's more convenient than dotTrace for preliminary perf work and is usually more than sufficient).
Also there is learning curve to grouping and aggregating data.
Yeah, that's never been true. Old compilers would often compile a switch to __slower__ code because they'd tend to always go to a jump table implementation.
A better reason to use the switch is because it's better style in C-like languages. Using an if statement for that sort of thing looks like Python; it makes the code harder to maintain.
(Also, Python has a switch-like construct now.)
PS: I'm presently revisiting C++14 because it's the most universal statically-compiled language to quickly answer interview problems. It would be unfair to impose Rust, Go, Elixir, or Haskell on an interviewer software engineer.
Very true, though there is one case where one can be highly confident that this is the case: code elimination.
You can't get any faster than not doing something in the first place.
Mostly the `final` keyword serves as a compile-time assertion. The compiler (sometimes linker) is perfectly capable of seeing that a class has no derived classes, but what `final` assures is that if you attempt to derive from such a class, you will raise a compile-time error.
This is similar to how `inline` works in practice -- rather than providing a useful hint to the compiler (though the compiler is free to treat it that way) it provides an assertion that if you do non-inlinable operations (e.g. non-tail recursion) then the compiler can flag that.
All of this is to say that `final` can speed up runtimes -- but it does so by forcing you to organize your code such that the guarantees apply. By using `final` classes, in places where dynamic dispatch can be reduced to static dispatch, you force the developer to not introduce patterns that would prevent static dispatch.
How? The compiler doesn't see the full program.
The linker I'm less sure about. If the class isn't guaranteed to be fully private wouldn't an optimizing linker have to be conservative in case you inject a derived class?
It is also an optimization hint, but AFAIK, modern compiler ignore it.
Need a way to make inlining heuristics ignore whether a function is inline https://gcc.gnu.org/bugzilla/show_bug.cgi?id=93008
(Bug saw a few updates recently, that's how I remembered.)
As a workaround, if you need the linkage aspect of the inline keyword, you currently have to write fake templates instead. Not great.
It's not just "you can have multiple definitions of the same function" but rather a promise that the function doesn't need to be address/pointer equivalent between translation units. This is arguably more important than inlining directly because it means the compiler can fully deduce how the function may be used without any LTO or other cross translation unit optimisation techniques.
Of course you could still technically expose a pointer to the function outside a TU but doing so would be obvious to the compiler and it can fall back to generating a strictly conformant version of the function. Otherwise however it can potentially deduce that some branches in said function are unreachable and eliminate them or otherwise specialise the code for the specific use cases in that TU. So it potentially opens up alternative optimisations even if there's still a function call and it's not inlined directly.
No, its purpose was and is still to specify a preference for inlining. The C++ standard itself says this:
> The inline specifier indicates to the implementation that inline substitution of the function body at the point of call is to be preferred to the usual function call mechanism.
Traditionally you'd use `static` for that use case, wouldn't you?
After all, `inline` can be ignored, `static` can't.
I can see exactly one use for an effect like that: static variables within the function.
Are there any other uses?
That's incorrect. The optimizer has to assume everything escapes the current optimization unit unless explicitly told otherwise. It needs explicit guarantees about the visibility to figure out the extent of the derivations allowed.
There are two applications, dynamic calls and dynamic casts.
Dynamic casts to final classes dont require to check the whole inheritance chain. Recently done this in styx [0]. The gain may appear marginal, e.g 3 or 4 dereferences saved but in programs based on OOP you can easily have *Billions* of dynamic casts saved.
[0]: https://gitlab.com/styx-lang/styx/-/commit/62c48e004d5485d4f....
For example, the AWS C++ SDK uses virtual functions for everything. When you subclass their classes, marking your classes as final allows the compiler to devirtualize your own calls to your own functions (GCC does this reliably).
I'm curious to understand better how clang is producing worse code in these cases. The code used for the blog post is a bit too complicated for me to look at, but I would love to see some microbenchmarks. My guess is that there is some kind of icache or code side problem. where inlining more produces worse code.
`final` tells the compiler that nothing extends this class. That means the compiler can theoretically do things like inlining class methods and eliminate virtual method calls (perhaps duplicating the method)?
However, it's quite possible that one of those optimizations makes the code bigger or misaligns things with the cache in unexpected ways. Sometimes, a method call can bet faster than inlining. Especially with hot loops.
All this being said, I'd expect final to offer very little benefit over PGO. Its main value is the constraint it imposes and not the optimization it might enable.
I want to ask, and I sincerely mean no snark, what is the point?
When working with AWS through an SDK your code will spend most of the time waiting on network calls.
What is the point of devirtualizing your function calls to save an indirection when you will be spending several orders of magnitude more time just waiting for the RPC to resolve?
It just doesn't seem like something even worth thinking about at all.
I don't think this really shows what `final` does, not to code generation, not to performance, not to the actual semantics of the program. There is no magic bullet - if putting `final` on every single class would always make it faster, it wouldn't be a keyword, it'd be a compiler optimization.
`final` does one specific thing: It tells a compiler that it can be sure that the given object is not going to have anything derive from it.
...and the compiler can optimize using that information.
(It could also do the same without the keyword, with LTO.)
"In theory" adding 'final' only gives a compiler more information, so should only result in same or faster code.
In practice, some optimizations improve performance for more expected or important cases (in the compiler writer's estimation), with worse outcomes in other less expected, less important cases. Without a clear understanding the when and how of these 'final' optimizations, it isn't clear without benchmarking after the fact, when to use it, or not.
That makes any given test much less helpful. Since all we know is 'final' was not helpful in this case. We have no basis to know how general these results are.
But it would be deeply strange if 'final' was generally unhelpful. Informationally it does only one purely helpful thing: reduce the number of linking/runtime contexts the compiler needs to worry about.
Why would I expect no performance difference? I haven't looked at the code, but I would expect that for each pixel, it iterates through an array/vector/list etc. of objects that implement some common interface, and calls one or more methods (probably something called intersectRay() or similar) on that interface. By design, that interface cannot be made final, and that's what counts. Whether the concrete derived classes are final or not makes no difference.
In order to make this a good test of "final", the pointer type of that container should be constrained to a concrete object type, like Sphere. Of course, this means the scene is limited to spheres.
The only case where final can make a difference, by devirtualising a call that couldn't otherwise be devirtualised, is when you hold a pointer to that type, and the object it points at was allocated "uncertainly", e.g., by the caller. (If the object was allocated in the same basic block where the method call later occurs, the compiler already knows its runtime type and will devirtualise the call anyway, even without "final".)
That definitely is one of the heuristics in MSVC++.
We have some performance critical code and at one point we noticed a slowdown of around ~4% in a couple of our performance tests. I investigated but the only change to that code base involved fixing up an error message (i.e. no logic difference and not even on the direct code path of the test as it would not hit that error).
Turns out that:
int some_func() {
if (bad)
throw std::exception("Error");
return some_int;
}
Inlined just fine, but after adding more text to the exception error message it no longer inlined, causing the slow-down.
You could either fix it with __forceinline or by moving the exception to a function call.std::exception does not take a string in its constructor, so most likely you used std::runtime_error. std::runtime_error has a pretty complex constructor if you pass into it a long string. If it's a small string then there's no issue because it stores its contents in an internal buffer, but if it's a longer string then it has to use a reference counting scheme to allow for its copy constructor to be noexcept.
That is why you can see different behavior if you use a long string versus a short string. You can also see vastly different codegen with plain std::string as well depending on whether you pass it a short string literal or a long string literal.
You're right, I used it as a short-hand for our internal exception function, forgetting that the std one does not take a string. Our error handling function is a simple static function that takes an std::string and throws a newly constructed object with that string as a field.
But yes, it could very well have been that the string surpassed the short string optimisation threshold or something similar. I did verify the assembly before and after and the function definitely inlined before and no longer inlined after. Moving the 'throw' (and, importantly, the string literal) into a separate function that was called from the same spot ensured it inlined again and the performance was back to normal.
The reason is placement new. It is legal (given that certain invariants are upheld) in C++ to say `new(this) DerivedClass`, and compilers must assume that each method could potentially have done this, changing the vtable pointer of the object.
The `final` keyword somewhat counteracts this, but even GCC still only opportunistically honors it - i.e. it inserts a check if the vtable is the expected value before calling the devirtualized function, falling back on the indirect call.
Kotlin (which uses the equivalent of the Java "final" keyword by default) uses the "open" keyword for that purpose.
Having said that "final" on member functions is great, and I like to see that instead of "override".
Changing an existing method way of calling (regular, virtual, static), changing visibility, overloading, introducing a name that clashes downstream, introducing a virtual destructor, making a data member non-copyable,...
C++ largely solves it by having tight encapsulation. As long as you don't change anything that breaks your existing interface, you should be good. And your interface is opt-in, including public members and virtual functions.
It doesn't go away just because private members exist as possible language feature.
Some APIs are aimed towards derived classes, like protected members and virtual functions, but that doesn't make the issue fundamentally different. It's just breaking APIs.
Point is, in C++ you have to opt-in to make these API surfaces, they are not the default.
But it basically boils down to uniform_real_distribution having a bunch of uninlined calls to 'logl' when compiled with Clang.
Otherwise, Clang beats GCC at least on the configuration I tested.
(I am the author of the issue)
[1] https://gitlab.com/define-private-public/PSRayTracing/-/issu...
Coincidentally, I happened to be playing around yesterday with a small performance test case using uniform_real_distribution, and for some strange reason Clang was 6x slower than GCC.
I put it down to some weird clang bug on my LTS version of Ubuntu. As my installed version was clang-14, I decided it possibly had been noticed and fixed a long time ago.
After reading your message I replaced uniform_real_distribution by uniform_int_distribution, and lo and behold, Clang was indeed faster than GCC, as expected.
Thank you for coming back to me with your findings.
[0]: https://research.facebook.com/publications/bolt-a-practical-...
In cases where you have Dog and Goose that both derive from Animal and then you have std::vector<Animal>, what is the compiler supposed to do?
As you say, that's the hot one -- and making the concrete subclasses themselves "final" enables no devirtualisations since there are no opportunities for it.
https://godbolt.org/z/7xKj6qTcj
edit: And a case involving inlining:
Also, now that I think of it, they should have run the code under perf and compared the stats.
During such long and compute-intensive tests, how are thermal considerations mitigated? Not saying that this was case here, but I can see how after saturating all cores for 8 hours, the whole PC might get hot to the point CPU starts throttling, so when you reboot to next OS or start another batch, overall performance could be a bit lower.
There will be operating system noise that can be in the multi-percent range. This is defined as various OS services that run "in the background" taking up cpu time, emptying cache lines (which may be most important), and flushing a few translate lookaside entries.
Once you recognize the variability from run to run, claiming "1%" becomes less credible. Depending on the noise level, of course.
Linux benchmarks like SPECcpu tend to be run in "single-user mode" meaning almost no background processes are running.
The clang regression might be explainable by final allowing some additional inlining and clang making an hash of it.
I am prepared to believe that there is some performance difference between the two, varying per case, but I would expect a few percent difference, not twice the run time..
See [1] for more information.
I started skimming this article after a while, because it seemed to be going into the weeds of performance comparison without ever backing up to look at what the change might be doing. Which meant that I couldn't tell if I was going to be looking at the usual random noise of performance testing or something real.
For `final`, I'd want to at least see if it changing the generated code by replacing indirect vtable calls with direct or inlined calls. It might be that the compiler is already figuring it out and the keyword isn't doing anything. It might be that the compiler is changing code, but the target address was already well-predicted and it's perturbing code layout enough that it gets slower (or faster). There could be something interesting here, but I can't tell without at least a little assembly output (or perhaps a relevant portion of some intermediate representation, not that I would know which one to look at).
If it's not changing anything, then perhaps there could be an interesting investigation into the variance of performance testing in this scenario. If it's changing something, then there could be an interesting investigation into when that makes things faster vs slower. As it is, I can't tell what I should be looking for.
It can't possibly be doing this, if the raytracing code is like any other raytracer I've ever seen -- since it must be looping through a list of concrete objects that implement some shared interface, calling intersectRay() on each one, and the existence of those derived concrete object types means that that shared interface can't be made final, and that's the only thing that would enable devirtualisation -- it makes no difference whether the concrete derived types themselves are final or not.
Fortran has virtual functions ("type bound procedures"), and supports a NON_OVERRIDABLE attribute on them that is basically "final". (FINAL exists but means something else.). But it also has a means for localizing the non-overridable property.
If a type bound procedure is declared in a module, and is PRIVATE, then overrides in subtypes ("extended derived types") work as usual for subtypes in the same module, but can't be affected by overrides that appear in other modules. This allows a compiler to notice when a type has no subtypes in the same module, and basically infer that it is non-overridable locally, and thus resolve calls at compilation time.
Or it would, if compilers implemented this feature correctly. It's not well described in the standard, and only half of the Fortran compilers in the wild actually support it. So like too many things in the Fortran world, it might be useful, but it's not portable.
In fact, I would run the same test repeatedly, keeping track of the k fastest times (k being ~3-7), and only stopping when the first and the kth fastest times are within a certain tolerance (as low as 1%). This ensures repeatability.
One sample of performance data for each test is not enough. This study provides no new insights.
Performance analyst
The best I'll see is somebody who cooked up a naive microbenchmark to show that style 1 takes fewer wall nanoseconds than style 2 on his laptop.
People I've worked with don't use profilers, claiming that they can't trust it. Really they just can't be bothered to run it and interpret the output.
The truth is, most of us don't write C++ because of performance; we write C++ because that's the language the code is written in.
The performance gained by different C++ techniques seldom matters, and when it does you have to measure. Profiler reports almost always surprise me the first few times -- your mental model of what's going on and what matters is probably wrong.
From a user perspective it could be the difference between software that's pleasant to use and software that's annoying to use. From a philosophical perspective it's the difference between software that functions vs software that works well.
Of course it depends on your context as to whether this is valued, but I wouldn't dismiss it. Once person's micro-optimization is another person's polish.
I'm disappointed the author's conclusion is "don't use final", not "something is wrong with clang".
Without a comparison of generated code, it could be anything.
In this way: you can avoid the need for the `final` keyword and do the optimization the keyword enables (de-virtualize calls).
>Yes, it is very hacky and I am disgusted by this myself. I would never do this in an actual product
Why? What's with the C++ community and their disgust for macros without any underlying reasoning? It reminds me of everyone blindly saying "Don't use goto; it creates spaghetti code".
Sure, if macros are overly used: it can be hard to read and maintain. But, for something simple like this, you shouldn't be thinking "I would never do this in an actual product".
(But I'm worse than the author; if I'm just comparing performance, I'd probably put `final` everywhere applicable and then do separate compiles with `-Dfinal=` and `-Dfinal=final`... I'd be making the assumption that it's something I either always or never want eventually, though.)
The main remaining use case for the old C macro facility I still see in new code is to support conditional compilation of architecture-specific code e.g. ARM vs x86 assembly routines or intrinsics.
Many people have a similar reaction to the use of "goto", even though it is absolutely the right choice in some contexts.
So, it's an opt-in security feature first, and a compiler hint second.
What actually may help is __attribute__((pure)) and __attribute__((const)), but I don't see them often in real code (unfortunately).
But you're right that this does not hold true for const pointers or references.
> What actually may help is __attribute__((pure)) and __attribute__((const)), but I don't see them often in real code (unfortunately).
It's disppointing that these haven't been standardized. I'd prefer different semantics though, e.g. something that allows things like memoization or other forms of caching that are technically side effects but where you still are ok with allowing the compiler to remove / reorder / eliminate calls.
https://godbolt.org/z/6ebrbaM7b
In general, whenever you call a function that the compiler cannot inspect (because it is in another TU) and the compiler cannot prove that that function doesn't have any reference to your variable it has to assume that the function might change your variable. Only passing a const reference won't help you here because it is legal to cast away constness and modify the variable unless the original variable was const.
I wish that const meant something on reference or pointers and you had to do something more explicit like a mutable member to allow modifying a variable. But even that would not help if the compiler can't prove that a non-const pointer hasn't escaped somehow. You could add __attribute__((pure)) to the function to help the compiler but that is a lot stricter so can't always be used.
Plus, you can’t even compile your code if you try to modify a const variable.
In the same TU, sure. But across TU boundaries the compiler really can't figure out what should be const and what should not, so `const` in parameter or return values allows the compiler to tell the human "You are attempting to make a modification to a value that some other TU put into RO memory.", or issue similar diagnostics.
Const can only ever possibly have a performance impact when used directly on variables. const pointers / references are purely for the benefit of the programmer - the compiler can assume nothing because the variable could be modified elsewhere or through another pointer/reference and const_cast is legal anyway unless the original variable was const.
Now for libraries, this is a different story. There I can imagine final keyword could have an impact.
I feel like we'd have to repeat these tests quite a few times to get to a decent conclusion. Hell small variations in performance could be caused by all sorts of things outside the actual program.
There are tons of these suggestions. Like always using sealed in C# or never use private in Java.
> I would never do this in an actual product
what, why?
It's is an insane level of ignorance about how these things are decided by the standards committee.
That's the first problem I see with the article. C++ isn't a fast language, as it is. There are far too many issues with e.g. aliasing rules, lack of proper vectorization (for the runtime arch), etc.
If you wish to have a relatively good performance for your code, try ISPC, which still allows you to get great performance with vectorization up to AVX-512, without turning to intrisics.
That's a bold statement due to the way it heavily contrasts with reality.
C++ is ever present in high performance benchmarks as either the highest performing language or second only to C. It's weird seeing someone claim with a straight face that "C++ isn't a fast language, as it is".
To make matters worse, you go on confusing what a programming language is, and confusing implementation details with language features. It's like claiming that C++ isn't a language for computational graphics just because no C++ standard dedicates a chapter to it.
Just like every engineering domain,you need to have deep knowledge on details to milk the last drop of performance improvements out of a program. Low-latency C++ is a testament of how the smallest details can be critical of performance. But you need to be completely detached from reality to claim that C++ isn't a fast language.
I'm ready to back this up. And no, I'm not confusing things - I work in HPC (realtime computer vision) and in reality the only thing we'd use C++ for is "glue", i.e. binding implementations of the actual algorithms implemented in other languages together.
Implementations could be e.g. in CUDA, ISPC, neural-inference via TensorRT, etc.
You a junior or something? For 99% of use cases C++ autovectorisation does plenty and will outperform the same code written in higher level languages. You are literally in the 1% and conflating your use case for that of the general case...
But to add to all the nonsense,you claim otherwise.
Frankly, your comments lack any credibility, which is confirmed by your lame appeal to authority.