HNHacker News
TopNewBestAskShowJobs

terrymah

161 karma · joined January 29, 2012

submissionscomments
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Until yesterday I thought I was the only person in the world who thought the designed behavior was undesirable, though
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Things are much better in 2024 in MSVC than they were in 2014. The overhead today is mostly the additional metadata associated with tracking the state, and most of the inline compatibilities were worked through (with a ton of work by the compiler devs). So it's a binary size issue. We've even been working on that (I remember doing work to combine adjacent identical regions, etc). Not sure what the status is in GCC/LLVM today.

I'm just a little sore about it because it was being sold as a "hey here is an optimization!" and it very much was not, at least from where I was sitting. I thought this was a very very good case of having it be UB (I think the entire class of user source annotations like this should be UB if the runtime behavior violates the user annotation)

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
We had a UB version of noexcept for a very, very long time. __declspec(nothrow), the throw() function specifier, etc.
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
No, calling throw in a noexcept function is a defined behavior (call std::terminate), and that behavior is not a diagnostic

I think maybe WG21 was concerned a compiler engineer would be clever if throwing in noexcept were UB, for example and assume any block that throws is unreachable and could just be removed along with all blocks it postdominates. Compiler guys love optimizations that just remove code. The fastest and smallest code is code that can’t run and doesn’t exist

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Determining if a function throws is a pretty basic bit of information collected in bottom up codegen (or during pre pass of a whole program optimization) and in no sense NP hard. Compilers have been doing it for decades and it’s useful

Noexcept on the surface is useful, except for the terminate guarantee, which requires a ton of work to avoid metadata size growth and hurts inlining. If violations of noexcept were UB and it was a pure optimization hint the world would be much better

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
I think WG21 has been violently against adding additional UB to the language, because of some hacker news articles a decade ago about people being alarmed at null pointer checks being elided or things happening that didn’t match their expectation in signed int overflow or whatever. Generally it seems a view of spread that compiler implementers view undefined behavior as a license to party, that we’re generally having too much fun, and are not to be trusted.

In reality undefined behavior is useful in the sense that (like this case) it allows us to not have to write code to consider and handle certain situations - code which may make all situations slower, or allows certain optimizations to exist which work 99% of the time.

Regarding “not pan out”: I think the overhead of noexcept for the single function call case is fine, and inlining is and has always been the issue.

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Everyone keeps scanning over the inlining issues, which I think are much larger

“Zero overhead” refers to the actual functions code gen; there are still tables and stuff that have to be updated

Our implementation of noexcept for the single function case I think is fine now. There is a single extra bit in the exception function info which is checked by the unwinder. Other than requiring exception info in cases where we otherwise wouldn’t

The inlining case has always been both more complicated and more of a problem. If your language feature inhibits inlining in any situation you have a real problem

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Well, yeah, things can be related to many things, but throwing extern "C"s was one of the motivations as I recall for 'r'. r is about a compiler optimization where we elide the runtime terminate check if we can statically "prove" a function can never throw. To prove it statically we depend on things like extern "C" functions not throwing, even though users can (and do) totally write that code.
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Nah that was mostly about extern "C" functions which technically can't throw (so the noexcept runtime stuff would be optimized out) but in practice there is a ton of code marked extern "C" which throws
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Dude I am going to blow your mind
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Oh, cool! I googled myself and someone actually archived the slides from the talk I gave. I think it holds up pretty well today

https://github.com/TriangleCppDevelopersGroup/TerryMahaffeyC...

*edit except the stuff about fastlink

*edit 2 also I have since added a heuristic bonus for the "inline" keyword because I could no longer stand the irony of "inline" not having anything to do with inlining

*edit 3 ok, also statements like "consider doing X if you have no security exposure" haven't held up well

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
No, we compile in bottom up order, starting with leaf functions, and collecting information about functions as we go. So "not throwing" sort of trickles up when possible to a certain degree.

In LTCG (MSVC)/O3 (GCC/Clang) there are prepasses over the entire callgraph to collect this order

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
It absolutely does, and even better, the compiler deduced "this function doesn't throw" doesn't come with the overhead of implementing noexcept proper
terrymah··on C++'s `noexcept` can sometimes help or hurt performance
You can't just look at the codegen of the function itself, you also have to consider the metadata, and the overhead of processing any metadata

Specifically here (as I said in other comments) where it goes from complicated/quality of implementation issue to "shit this is complicated" is when you consider inlining. If noexcept inhibits inlining in any conceivable circumstances then it's having a dramatic (slightly indirect) impact on performance

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
In MSVC we've also pretty heavily optimized the whole function case such that we no longer have a literal try/catch block around it (I think there is a single bit in our per function unwind info that the unwinder checks and kills the program if it encounters while unwinding). One extra branch but no increase in the unwind metadata size

The inlining case was always the hard problem to solve though

terrymah··on C++'s `noexcept` can sometimes help or hurt performance
Oh man, don't get me started. This was a point in a talk I gave years ago called "Please Please Help the Compiler" (what I thought was a clever cut at the conventional wisdom at the time of "Don't Try to Help the Compiler")

I work on MSVC backend. I argued pretty strenuously at the time that noexcept was costly and being marketed incorrectly. Perhaps the costs are worth it, but none the less there is a cost

The reason is simple: there is a guarantee here that noexcept functions don't throw. std::terminate has to be called. That has to be implemented. There is some cost to that - conceptually every noexcept function (or worse, every call to a noexcept function) is surrounded by a giant try/catch(...) block.

Yes there are optimizations here. But it's still not free

Less obvious; how does inlining work? What happens if you inline a noexcept function into a function that allows exceptions? Do we now have "regions" of noexceptness inside that function (answer: yes). How do you implement that? Again, this is implementable, but this is even harder than the whole function case, and a naive/early implementation might prohibit inlining across degrees of noexcept-ness to be correct/as-if. And guess what, this is what early versions of MSVC did, and this was our biggest problem: a problem which grew release after release as noexcept permeated the standard library.

Anyway. My point is, we need more backend compiler engineers on WG21 and not just front end, library, and language lawyer guys.

I argued then that if instead noexcept violations were undefined, we could ignore all this, and instead just treat it as the pure optimization it was being marketed as (ie, help prove a region can't throw, so we can elide entire try/catch blocks etc). The reaction to my suggestion was not positive.

terrymah··on Endless Doom Scroller
I, too, came for the procedurally generated Doom levels and left disappointed
terrymah··on Layoffs Are Coming
From direct effects, maybe. But no one is immune from the second order effects. Once our customers start going out of business because they are directly exposed, they can't buy our software anymore.
terrymah··on History of Xenix – Microsoft's Forgotten Unix-Based Operating System
It was a joint project, with ms eventually pulling out to work on NT
terrymah··on What is the Nash Equilibrium and why does it matter?
You sort of have to, since everyone else is doing it.
terrymah··on Introducing a new, advanced Visual C++ code optimizer
This is both C and C++ - they share a common backend (c2.dll) where this work was done. They have different front ends (c1.dll / c1xx.dll).
terrymah··on Why Registers Are Fast and RAM Is Slow
Hmph. You forgot the part where Bubba's friends are watching him drink beer and eat Hungry-Mans, and if they want some, they can force Bubba to throw out all his food and pour all his beer down the drain, and everyone has to go back to the store.
terrymah··on Why Registers Are Fast and RAM Is Slow
I've heard that modern Intel processors have 100 < x < 200 physical registers. I'm not sure they actually document the exact number.
terrymah··on Why Registers Are Fast and RAM Is Slow
It's complicated, but modern processors actually do have many more registers than you can name in the instructions. They use things like "register renaming" to avoid false conflicts between instructions.

Registers that you name in assembly != physical registers. And when you use a register in two different instructions, you won't necessarily get the same physical register each time.

terrymah··on Humor: Interview with an Ex-Microsoftie Who Used to Name OS Folders
FWIW, "WOW" stands for "Windows on Windows"
terrymah··on Visual Studio 2013
> PGO compiles only 0.4% of our application for speed, and our response times are about 20 usecs slower with it on

This is a complex issue, but consider abandoning PGO and just compiling for for speed then. PGO doesn't help in each and every case.

> RE xperf (and WinDbg): Stop bundling this shit in "Toolkits".

How else would you bundle it? It's not simple to put something as part of the base OS image. And I haven't heard about these installers breaking the VS installer - that sounds like a bad bug.

> __assume is so useless. How often does someone write a branch that does nothing every time?

It's commonly used as a retail version of a debug ASSERT macro. But yes, like I said earlier - I wish we would do more with static annotations, but I've gotten push back.

> PogoAutoSweep crashes threaded programs if you don't suspend every other thread but it's still quasi documented. The PogoSafeMode build flag/environment appears to be ignored.

I've never seen PogoAutoSweep crash - do you have a repro? PogoSafeMode doesn't affect PogoAutoSweep, only probe generation.

> The filename postfix that PogoAutoSweep adds breaks the VS2012 PGO menu options.

Haven't heard of this either, but stay tuned. I don't like the PGO menu options as they currently stand.

> There's nothing one can do to limit the VS2012 profiler to specific threads.

I can forward that request to the profiler team.

> The interface for instrumenting specific functions is terrible, use a plain text file or decl_spec FFS.

Are you talking about PGI or an instrumented profiler?

> If there are #defines or other ways to detect an instrumented build, they're terribly documented.

There isn't an easy way, and having different code in the PGI build versus the PGU build would be problematic.

> PGO instrumentation/optimization is woefully obtuse. What did it pick for speed? Why did it pick it? What branches did it fold/unfold?

Stay tuned

> How does the pgc weighting actually work?

The obvious way, the counts are multiplied by the provided factor before being merged in the PGD.

> Can I artificially create my own pgc?

Not realistically.

> Not related to our main response loop, but we can see in our logging threads that the LFH malloc appears to often call RtlAnsiStringToUnicodestring.

No idea (CRT owns malloc, Windows owns LFH).

> Speaking of which, changing the malloc implementation is still horrible even after the VS2010 msvcrt changes. In linux...

I'm not an expert, but my understanding was that malloc and friends were weak symbols, and if you just linked in an obj that defined malloc it would be selected as the "real" malloc without giving an ODR.

> Why is there SemaphoreSlim in C# but not C++? Why is there no Benaphore primitive that can also be used in WaitForMultipleObjects?

I'm not sure, Windows owns this.

> Serious issues in Microsoft Developer Connect are often ignored, closed as behaves as expected, or dismissed off hand

I've heard complaints about MSConnect before as well. All I can say is that it is the correct place to file bugs; and the issues there do directly show up in our bug list (someone goes through connect issues, filters/combines them, and files bugs).

> There is still no valgrind/cachegrind equivalent that provides the same level of detail

That is correct. Sorry.

> Our statically linked application takes 20 minutes link and the link is not parallel. C++ compiles are likewise brutally slow

Link.exe performance is at the top of our minds right now, you're not the only one to bring it up. VS 2013 will have some perf improvements across the FE (to help with C++ being brutally slow) but there is always more to do.

> And no, I'm not going to turn on precompiled headers, MSVC builds incorrect binaries about 5% of the time as it is.

Never heard that before - codegen bugs are always deadly serious and treated with high priority. If you have a repro, please share it.

terrymah··on Visual Studio 2013
Following the ABI is only an issue at module boundaries. If you control every callsite of a function, you can invent whatever calling convention you want.
terrymah··on Visual Studio 2013
Visual C++ Team blog
terrymah··on Visual Studio 2013
I'm actually about to sit down and write a blog post about VS compiler memory issues, and I had a long talk with the Firefox guys a few weeks ago about this issue. It's not all their fault.
terrymah··on Visual Studio 2013
Current VS-compiler dev here, the backend codegen team to be more precise (I actually own PGO). I wish this comment had been written in few days/weeks/months so I answer directly and talk specifics about some of the work that went into VS2013, but for now I just want you to know we've aware of all of the issues you brought up, and have either worked on or plan to work on many of them.

RE: likely/unlikely, VS has __assume(0), which isn't exactly the same thing I know, but it is something and does help. I'm actually in favor of us doing more with static annotations to bring PGO style optimizations to non-PGO builds. If you feel the same way please be louder about it, but realize there is a vocal group of people who consider static annotations harmful (and they have a large body of evidence in __forceinline backing them up).

oprofile: There is ETW/xperf, and of course a variety of instrumented profilers (both shipping and internal)

Although I do wish my team was larger, and it doesn't get all the love that some of the more flashing UI stuff does, I wouldn't go as far as to say the toolchain is withering. Some of the smartest people I know are working on my team with me on these problems.