C++ exception performance three years later
databasearchitects.blogspot.com
databasearchitects.blogspot.com
I wrote a program that throws an exception and catches it to warm up, and then throws a second exception.
void first() {
throw "";
}
void second() {
throw 2;
}
int main(int argc, char *argv[]) {
try {
first();
} catch (...) {
}
second();
}
The second() function with LLVM libcxx using normal DWARF exceptions calls pthread_rwlock_rdlock SIX TIMES and calls pthread_rwlock_wrlock one time. Overall, this program triggers ~4000 function calls according to --ftrace (see https://justine.lol/ftrace/)On the other hand, if I use SJLJ, then this program triggers 269 function calls. It never locks anything. Although, like the first program, it does call pthread_once three times. That's doesn't cost anything though. It's the rwlocks you need to worry about.
For an example of a modern C++ toolchain that uses SJLJ exceptions, see https://github.com/jart/cosmopolitan/releases/tag/3.7.1
Please note that SJLJ exceptions carries a tradeoff. While SJLJ will make throwing exceptions much faster, you'll incur a slight performance degradation on normal code that isn't throwing exceptions. Using GCC in SJLJ mode will cause most of your C++ functions to have extra code inserted into the prologue and epilogue which calls _Unwind_SjLj_Register and _Unwind_SjLj_Unregister. Modern DWARF exception handling makes the intentional tradeoff of making exception throwing extremely slow, so that zero overhead is added for normal non-exceptional code. Both methods (SJLJ and DWARF) will bloat your binary size. With SJLJ it's just a few instructions inserted into prologues and epilogues. But with DWARF you get a lot of mandatory .eh_frame junk that's generated by everything, and it's centralized into a particular section of your binary that's only touched if exceptions are actually thrown.
Warm -> populated with the code/data your program uses. Cold -> empty, or filled with stuff your program doesn't use.
I do think that the case of exception handling is one where I would question the validity of this approach: if you catch exceptions rarely, it's quite reasonable to expect that the relevant lines would have been conflict'ed out of your l1i at some point in the interim, and therefore that the l1i (or even l2!) fetch cost is reasonable to include in the benchmark. However, if you expect your exceptions to fire 'relatively often' then it again becomes reasonable to warm up the caches before measuring, so it really comes down to the characteristics of your workload.
Since you have only shown a simple example of throwing exceptions I'm not sure I understood your point how SJLJ exceptions solve the scalability problem? Truth to be told I'm not very familiar with SJLJ.
DWARF exception handling needs DWARF unwind information, which is mapped in memory along with the code. But the stack only has program counters, not the addresses of the DWARF data. Therefore, the first step is to locate the DWARF data. The semi-portable interface (supported on GNU/Linux, various BSD, and likely others) for that is dl_iterate_phdr, and that invokes a callback under a global lock. The global lock is required because the list of objects can change in response to dlopen and dlclose calls. With the new way, the GCC unwinder directly asks the glibc dynamic linker for the ELF object corresponding to a program counter value and its DWARF unwind information address. This way, the synchronization with dlopen/dlclose is an internal dynamic linker implementation detail.
https://wg21.link/p2544#expected
>> the overhead is so high that std::expected is not a good general purpose replacement for traditional exceptions.
I'm guessing you don't care about performance that much.
It seems pretty correct to me? It's just based on first principles. Returning a value + error is strictly more work than returning just a value, and compilers can't solve the halting problem to optimize away arbitrary amounts of extra work. Ergo, performance hit.
I'm comparing in-band to out-of-band error signaling, not "error handling to no error handling".
My understanding is that Zig errors can't contain any data. But an error message like "file not found" is awful. Are your error messages awful, or is there some way to get information into them so that you can say "file 'foo/bar.json' not found"?
If you need more than that and error codes, then either log as the sibling suggested or use custom return types. Anything more complex than what can be encoded as error codes is not an exception.
I think the issue is that many languages lacked support for option types for a long time so exceptions got overused and abused. Now the pendulum has swung towards "exceptions are bad".
I think both ways of handling errors serve different use cases and it is good to have proper language support for both. I like to use exceptions for errors that are... literally exceptions. Stuff that is out of my control, being out of memory, stuff that I don't expect to happen on the happy path. When the default should be to just log the thing and crash. And if your language supports proper exceptions you get nice stack traces in your logs to see what went wrong.
On the other hand error types are superior for things that you should explicitly handle in your code. That are reasonable invariants to take care of. Here the language forcing you to handle them or to explicitly opt out of handling them is extremely valuable.
Having more tools is not always bad.
[1] https://github.com/Microsoft/DirectXTK/wiki/throwIfFailed
> The root cause is that the unwinder grabs a global mutex to protect the unwinding tables from concurrent changes from shared libraries.
> We did a prototype implementation where we changed the gcc exception logic to register all unwinding tables in a b-tree with optimistic lock coupling.
Did they try just stripping synchronization altogether to see what the performance ceiling would be?
In the late 90s/early 2000s when languages started promoting general purpose threading, they did so without planning and simply threw locks on everything.
Everyone now pays the cost of that feature whether you use it or not.
Is throwing and catching exceptions across threads a good idea?
What is the state of current compilers for the non-throwing case? If I have code with exceptions, but test a case that does not actually throw? Will it be as fast as the same code but with exceptions removed altogether?
EDIT: I just saw that Justine's (jart) comment answers this.
Always returning a slightly "fatter" type must certainly have an associated cost as well and it is to pay in any case - error or not.
At the limit limit, there is no difference between Result<T,E> and checked exceptions, with enough syntactic sugar they can be equivalent both semantically and syntactically [1].
[1] One advantage of typical C++ exceptions implementations, is that if the exception is not caught the stack is not unwound and the stack trace is preserved in the core file. In principle the same thing could be done for Result.
So the real question is syntax - which do you prefer: having to write some result type all over instead of the regular return type you mean, or knowing that "magically" at unpredictable times your function will just abort in the middle. There are pros and cons to both, so anyone arguing absolutely for either side is wrong - at least now when we don't have decades of experience with result in large projects (though maybe we do in a single project)
Once you convert them in CPS form there is really no difference between the two. In principle you could compile both to exactly the same code. [edit: also not sure why are you singling out checked exceptions]
[edit: specifically you could compile any code that returns an Either to a function that takes two return continuations instead of just one, and in CPS form exception-like non-local exits are trivial]
Of course in practice traditional procedural languages do not compile through CPS and use abnormal edges in basic blocks to represent exceptions.
> There are pros and cons to both, so anyone arguing absolutely for either side is wrong
I'm in the camp that, in a well designed language[1], there would be no difference between the two :D
[1] according to my personal preferences that I can assure you are completely objective!
How do you figure? A Result<T,E> always means a branch after function calls, while checked exceptions have no such cost.
Toy example: https://godbolt.org/z/Gqo9n17c6
It's trivially observed that exceptions simply have less code in the normal path. There's no post-call checking for errors, which is of course present in the Result<T,E> approach. And even where the exception is caught, that's not part of the "hot loop". It's not executed at all normally. Unwinding exceptions are completely free when they don't happen. The same is simply not true and never can be true of a return-value based propagation.
It isn't easy, though, as the Rust way of doing this--letting people pass around a bare Result without the structure of a monad (which is why Haskell's version of this idea looks like exceptions, and why it is so frustrating when people incorrectly claim Rust has "monadic" error handling)--makes it extremely easy to pull off what should be "stunts", such as holding onto a Result and trying to resolve them in an irreducible order; but, "hopefully", most code isn't making that mistake (and you can still implement it by adding a spurious expensive catch in just the places they did that)... but like, that's why I think the person you are responding to might have said "at the limit limit", to really emphasize that this is a pretty complicated reduction to automatically apply globally and truly result in 0 cost.
One idea is merely a form of control flow (like 'if' or 'goto') that could jump down the call stack. This kind of exception handler could be made fast if necessary.
The other idea is to trigger the processor's actual fault handler, and proceed to the operating system, which then dispatches the exception to the program. The advantage of the latter is that you can catch exceptions raised by the processor that way. But since you need to enter the operating system's exception handler code, it goes slower. C# exceptions are like this.
Interestingly this is more or less how Symbian handled exceptions back in the day. It was a giant mess. You basically called setjmp in a surrounding function (using a macro IIRC) and then longjmp (in the form of user::leave()) which unwound the stack without calling any destructors.
It was very fast and the source of endless memory leaks.
It was also before RAII was a thing and so you had to manage the ‘CleanupStack’ [0] yourself.
At least for the development of the OS itself, and not frameworks such as the awful Series60 UI, we had have unit tests that brute forced memory correctness by getting the memory allocator to deliberately fail at each next step of the code being developed until it succeeded or panicked.
[0] http://devlib.symbian.slions.net/s3/GUID-E7D29464-05E1-5039-...
That speedup isn’t small. The largest one is for 8 threads, 0.1% vs 0%: 32 vs 64, so occasionally throwing exceptions seems to double the speed of that micro-benchmark.
Is there so much noise in that data? If not, what could cause this?
Exceptions are not meant to replace error codes, that you use to get a detailed picture of the current environment. They are rather like a kind of assertions, where you don't actually care too much what happens next, because the exception being thrown means someone, somewhere, somehow, has already fuc*ed up.
I am sure they have very good reasons to keep working on this issue. That's why I started my comment with "in general".
Just to state the obvious, this is a design philosophy (one of many) and not something imposed by the language. And reasonable people can disagree about whether it's a good one.
I think where a lot of engineers go wrong is where they go crazy baking exception-throwing into their logic and error handling code, leaving a beautiful "happy path" in their code, but they are not disciplined about actually catching and handling the exceptions. If you're going to go all in on throwing, then you need to go all in on catching. Otherwise your product is just going to be full of unhandled crashing. A couple projects worth of this, and a developer will start to think "hmm, exceptions are bad, I agree we shouldn't use them!" Which is how we get blanket company policies against using any exceptions at all.
However, in this case this abuse provided a nice boost for everyone, so I can't complain.
Excerpts from paper:
> Single threaded the fib code using std::expected is more than four times slower than using traditional exceptions.
> This has much less overhead than std::expected, but it is still not for free. For fib we see a slowdown of approx. 60% compared to traditional exceptions, which is still problematic.
Now, admittedly, these workloads are too simple (sqrt and fibonacci) and only shown for the single-threaded use-cases but I guess you have to start from something that is easy enough to reason about otherwise getting to any sane conclusions might be too difficult.
However, they continue to show in their blog how they significantly sped up the std::exception implementation in a very non-trivial workload - multi-threaded and JIT'ed database kernel.
So, to prove the point about exceptions being slower in complex workloads, such as the one above, somebody would actually have to rewrite their whole codebase to use something else such as std::expected, return values or whatever. Non-trivial to say the least.
[1] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2022/p25...
"Unwinding now scales with the number of cores, and we can safely use C++ exceptions even on machines with large core counts."
Basically, a naive implementation of exceptions uses a lock to guard unwind tables, and this scales poorly. The lock is needed only because .so libraries can be loaded and unloaded dynamically. One solution is to say that after a certain moment, the program doesn't load/unload dynamic libraries and doesn't use the lock after that. That's what ScyllaDB and Yandex[0] do, for example. Another solution is to just use the new glibc which fixes this issue.
That aside, both exceptions and error codes introduce overhead, but they do it differently: exceptions add most[1] of their overhead when an error happens, while checking error codes adds overhead when an error doesn't happen[2]. One might conclude that exceptions should be used when errors are exceptionally rare; otherwise, use error codes. Or, just reclassify errors that happen too often as something you expect to happen, return them as a case of valid data and keep using exceptions.
[0] https://habr.com/en/companies/yandex/articles/852244/
Many early programming languages, including various dialects of Algol and Fortran, included a better implementation method.
The invoking procedure passed to the invoked procedure not a single return address, but multiple return addresses. One return address was for the normal return and the other addresses were for alternate returns when errors were detected.
On a modern CPU with many registers, all these return addresses would be passed in registers.
The invoking procedure had at its end one or more handlers for error conditions, corresponding to the alternate return addresses. When a procedure was invoked, no tests were done at the invocation place, because any error would cause a jump to the error handler. In the invoked procedure, a test and a conditional jump must always exist in order to detect an error, and there the alternate return was executed only when an error was detected, so this implementation could omit one test and one conditional jump in comparison with the modern implementations.
Moreover, the text of the procedure was more clear in this way, with all the error handlers separated from the normal execution case, but nonetheless close enough to examine them when necessary.
I remember a case where we were using a library throwing exceptions when it could not parse some string. I think we were parsing IP addresses, or something like this. There was a case where this was called in a hot loop (like thousands of IP addresses in a roe) and for whatever reason in some cases most of them were invalid, thus throwing. Then simple error handling: the caller catches the exception, logs the error and checks the next one.
The program was frozen for seconds. Replacing the parser with the same logic but returning a bool instead was something like 1000x faster. That was on MacOs Clang Arm.
I don't see how a few extra ifs (though none were required in this case) would change anything.
Generally I'm not a fan of exceptions, you never what something is going to throw. The majority of the crashes we see in production are unhandled exceptions. Turned out the documentation did not mention that this function could also throw std exceptions on top of the lib ones in this specific case.
But, the other extreme you get is saying all error handling should be done with checks and branches, and the idea is that this slows down the code which isn't throwing lots of exceptions so much that, for average workloads, it is the wrong tradeoff.
Ideally, you could make this decision more locally, and yet continue to use the cleaner syntax which comes from structured error handling, but I don't know of any language which implements this :/.
And, so, the best we have is to just use a language which maps exception syntax to tables as you can always -- as you did -- fix the few places it matters to use a result type by adding manual checks.
"Make sure your performance-critical code / tight-loops don't do things which can throw exceptions, and then you won't have to worry (much) about the performance of exceptions"
?
It may not be what the article says, but it's how I avoid having to delve into exception performance.
Consider for example out of memory. "proper" error handling via return values would mean everything that allocates memory has to have an error return path and everything that calls anything that has that somewhere in the chain needs to still propagate that error. Essentially every function call now has a branch after it. If done explicitly, that's a lot of code nobody wants to write and, more importantly, that nobody ever actually tests. If done automatically, that's still just a metric shitload of branches which wastes branch predictor entries, it wastes cache size, etc...
Exceptions, meanwhile, are typically free unless actually thrown. So for the truly 0.001% situation, like the aforementioned out of memory, they are perfect. You also get stack traces for free which greatly aid in debugging.
Also typically you either handle errors pretty close to where the occur, or relatively high up at a more macro level. Exceptions lets your "middle" layer be significantly less branch-y & verbose for situations where it's a more macro-level error handling strategy.
The alternative approach, like what Rust does, is to just say that rare errors like that are unrecoverable and hard-crash when encountered. That's certainly a choice you could make, but it's pretty heavy-handed for the language itself to mandate that. Which is then why Rust also provides std::catch_unwind, but with notes that it doesn't necessarily work because people can have chosen to compile with panic set to hard-abort period. It's pretty much a big mess.
That seems like a win to me for the aforementioned "readable" criteria—failures should be enumerated. But I can see how that would be annoying to write.
I'm also of the opinion that if handling out of memory is a serious consideration for your process, you should consider using your own allocator that doesn't fire an exception. This is how rust handles it, for instance. And on top of this I'm a little confused because exceptions are mostly disallowed from embedded codebases, which is the primary place where you would likely run out of memory and want to actually handle it.
Sure they do - they don't negatively impact the non-error case and don't pollute every function's return value.
Handling some errors in band is simply not useful in modern programs. While kernels and embedded use cases benefit from in band allocation failures, most applications would do better to receive an out of band signal that they are close to memory thresholds and to perhaps shed load by doing things like stop accepting new sockets connections, shrink some caches, etc. I really used to think otherwise but too much software has been written to assume they will never experience a memory failure in band.
Also to clarify, I don't think an exception is out of band. They have a significant cost to program size and performance even if they are never used.
They do have a cost to program size, yes, but they have ~zero impact on performance when never used. They are completely cold, outside of normal flow.
Exceptions don't eliminate error checking conditionals, they just hide them elsewhere in the code.
The only thing I could see is just the density of code could mean more page spills and less efficient cache line usage for code. There's otherwise no actual cost to the assembly executed itself. And that's compared to no error handling at all, even. Compared to return values it's very clearly faster. But maybe there's some edge case you're aware of that does something different.
Some objects cannot be safely destroyed at arbitrary points during program execution because it will corrupt memory, leak resources, etc. In a lot of systems-y code you can have things in-flight that the code can’t reliably stop or undo as an intrinsic property of the system. If errors occur, you have to partially defer the unwinding of state indefinitely (or block) until it is safe to do so. This is pretty idiomatic for things like high-performance I/O.
Making this exception safe requires much more defensive programming. You need extra code on the happy path to collect context in case of an exception, extra code in the error path to interpret that context to figure out what is and is not safe to destroy at the point in execution where the exception occurred, and extra code in the handling path to execute the deferral logic based on an interpretation of the context.
With return code style handling, this is usually implicit in the structure of the code. It is much nicer to do this at compile-time than run-time. If you only write high-level C++ then this might not apply but most C++ tends to be used for systems-y software these days. The widespread practice of disabling exceptions for C++ code bases at companies big and small didn’t happen for no reason.
Again, provide an actual example please. C++ exception context handling is generated at compile time and is in a secondary, cold stack.
It's unclear if you actually know of something specific or if you're just speculating on what you assume happens instead of what actually does