Cost of a thread in C++ under Linux
lemire.me
lemire.me
MacOS doesn't have this problem but Linux and FreeBSD do.
Edit: last time I investigated the issue ended up here: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=71744
If my caller has destructors to run or a catch clause this pointer is null and inspection proceeds as normal. If It does not it stores the value from its frame there. Then if I throw an exception I jump to the next frame that needs inspection; if I don’t then any throw further down the call stack won’t even look at me.
The C++ standard can’t call for this because of the “zero cost if you don’t use it” rule. But a Linux ABI could. The MacOS takes advantage of this kind of freedom.
> The C++ standard can’t call for this because of the “zero cost if you don’t use it” rule. But a Linux ABI could. The MacOS takes advantage of this kind of freedom.
That's not really true. It has nothing to do with the standard. It has everything to do with the compiler's users complaining about the performance hit relative to DWARF EH. It is part of the social contract between the standards body, the compiler author community, and the user community that unused features don't cost us in runtime performance.
As for "policy" vs "social contract" I think we basically agree.
And when you are unwinding a lisp->C++ boundary (that is, lisp code called by a C++ function) you are free to do what you like until you get to the first Lisp frame; if it doesn't have an unwind-protect then its "ignore me" pointer just points up to its caller, which is examined by the C++ runtime anyway.
The nice part of that second paragraph is that if that first lisp callee was called by a non-c++ function (say a fortran function) you might even have an opportunity to set the "parent frame for inspection pointer" to skip over all the fortran frames and point directly to the lowest C++ function below you...which you could manage via a small change to gold or llvm-ld.
Apple does indeed have more freedom, and it may be that specific MacOS components need this in ways that the general community doesn't seem to. But I'd want to see numbers from a bunch of real world environments before declaring this a uniformly good optimization.
I'm glad if it is so. Exceptions should actually be "exceptional" and not the part of the normal execution flow. Whoever has other ideas has the wrong model of what, at the lower levels, exceptions actually do.
It is distressing that Sutter's survey showed that half the respondents had to disable exceptions for part of all the code. I've often heard the argument "well google's coding standard prohibits exceptions" which is bizarre, as google's standard says "exceptions are great but we have some legacy code that can't use them, so we're stuck"
The biggest argument seems to be that they are expensive, which is crazy because there's no cost if you don't raise one and if you do you're already in trouble and generally have plenty of time to deal with it (this is different from, say, Lisp signalling which not only permits continuing (!) but is on theory supposed to be common. Probably a mistake in retrospect). But they allow you to make the uncommon stuff uncommon (as opposed to error codes which must be sprayed like shrapnel through your code).
There are two legit arguments against exceptions: one is when you are constrained in space (e.g. embedded systems) and/or time (hard realtime systems that need predictable timing, even if it is slower). The other is a philosophical argument that it embodies a second, parallel flow of control. Since C++'s exception system is an error system only, and since destructors are run automatically, it's hard for me to find this second argument convincing.
That is not the C++ I enjoy using, rather the language I got to love via Turbo Vision, OWL, VCL, MFC, Qt, which is not what drives the language nowadays.
Take a look At C++ (or c++ 20!) as if it were a brand new language you’d never seen before and forgetting that it’s name includes “c”. That language is a pretty clean, expressive and straightforward language IMHO. I like programming in it.
It’s not claiming it’s unicorns farting rainbows, but it’s definitely pretty good.
Yeah, if everyone plays ball, it might come in 5 years from now, assuming C++23 gets done on time, plus the compiler support stabilization.
Right now SG14 seems to drive some of those decisions, at least from outside.
The standard can move quickly: consider formatted output which lingered unchanged with a broken model but was rapidly reformed when someone with a good model and implementation was encouraged to come forward. Admittedly a smaller topic than concurrency or networking!
The other language communities manage to drive language progress over the Internet, which apparently ISO has yet to get in touch how it goes.
The one doing application stuff in Qt, MFC, wxWidgets.
Or those like myself, where C++ only matters as means to implement native bindings to system libraries, or GPGPU shading languages based on C++.
So I would benefit from any number of abi-breaking proposals but can understand the committees reticence.
Also I believe that ABI compatibility with binaries compiled with different compilers also is complicated by things like name mangling conventions which are not common between compilers.
There are so many things that can go wrong that the people i know that rely on ABI compatibility in C++ make a diff of their objdump output a ci test. Sometimes they get more paranoid and send me an assembly diff when they are suspicious(a compiler upgeade generated different jump prolog:) ). Ridiculous compared to the price of just running the compile again. The fact we control the whole machine's os image even makes it harder to understand... Oh man the pain I have...
My apologies if my text is weird but proof reading in a phone is hard.
[1]https://community.kde.org/Policies/Binary_Compatibility_Issu...
Unfortunately some people do use downvoting to mean "I disagree", presumably based on the convention of some other forum. There's not much to be done about that except upvoting posts you see have been unfairly downvoted.
This is just false in the general case. The presence (potential or actual) of exceptions often just serves as an optimization barrier in current compilers. That's not to even invoke bizarre but not infrequent issues like this [1]. I too have had codebases that miraculously sped up upon disabling exceptions despite not throwing anything. Identifying the exact causes of these situations is hard and typically not done, because it's far easier to just add a compiler switch and pretend there are no exceptions in C++ and get back to work.
Many people in performance sensitive domains just don't find it remotely worthwhile to care about features that have these sorts of difficult to predict and debug costs. When your workflow already consists of writing highly explicit, simple to reason about code that you frequently inspect in disassembled form, exceptions (and RTTI for a host of obvious reasons) are the last thing you'd want to enable. At best it's just extraneous noise in the assembly, at worst you take a sizable perf hit and have no idea why.
[1] https://twitter.com/timsweeneyepic/status/122307740466037145...
A workflow consisting of writing highly explicit, simple to reason about code that you frequently inspect in disassembled form
Is possibly the most descriptive and succinct description of my coding practice. I really like this formulation and I am going to shamelessly steal it in the future, repeatedly.
As far as what I’d consider frequent use for me, the only feature I use heavily is namespacing (and that is really only for personal organizational benefits). I do take advantage of standard templated containers and classes when I think they are the right call. Generally, the use of classes and associated method mechanisms are really only used for specific circumstances (which are purely aesthetic for me), and I certainly lean toward custom containers if I can. I do use operator overloading for math, but that’s about it. It is pretty rare for me to use any inheritance, virtual functions, etc, and I don’t think I have ever programmed any exceptions, but maybe some libraries have them, same for RTTI. I do use BLAS, and some other template based libraries.
I don’t use much from after C++98, but I do occasionally dip into C++11 for constexpr. I don’t ever use auto or decltype and I may have looked at range-based fors, but they aren’t used anywhere I can recall.
I think if I could have actual namespaces, instead of space_variable style, C would do it for me. I certainly like that restrict is part of the language, not just a compiler intrinsic. But, and it is a big but, while Clang/llvm, GCC (for the most part) and some proprietary C compilers are good enough for my purposes, MS’s C compiler is barely mediocre as far as I’ve heard (I just went with common talking points when I decided to go down the C++ route, and haven’t actually tested equivalent implementations). Also, while making shim layers is possible, C++ libraries are widespread and common, and just using C++ is less friction and maintenance.
If it isn’t obvious from all of that, my coding style in C++ is basically a rip-off if Mike Acton’s CPPCon talk in 2014. If I want anything more abstract for some reason, I’ll use Python or Haskell or Racket (and recently Ocaml).
Well, don't. "Exception-based programming" is an anti-pattern. Exceptions should be thrown in, well, exceptional situations.
[1] https://eli.thegreenplace.net/2018/launching-linux-threads-a...
[2] https://eli.thegreenplace.net/2018/measuring-context-switchi...
Some other things: pthreads generally have high cost, and that means C++ threads do too. pthreads have quite a few features that you regularly don't use, which you can skip completely using fibers and coroutines.
I don’t think this is a realistic expectation on Linux and, especially, Windows which runs hundreds of threads of its own you don’t want to know about. (Besides, we must remember that multithreading was invented and found quite useful in the era of ”single-core” processors.)
Edit: to clarify: any writeable mapping or any unmap will cause TLB shutdown interrupts to be broadcasted to all currently running threads of a process.
These pauses happen even without unmapping. Despite that no synchronization is available, so you are right in the middle of whatever, the kernel decides a static snapshot of the pages' state must be written, so write-protects the pages first, and blocks your process until the write is done. It's just rude.
yes, that sort of spikes are definitely not due to just interrupts. As you say, it is probably a kernel io thread taking over the cpu.
If we're talking 50%, the complexity sounds worth it, but if it's 1% I think I'd prefer to stick with standard scheduling and know my program will 'just work' on any CPU or OS, and with any libraries I choose to use.
The reason why is because we avoid all the indirect cost of context switching, which is all the various caches that has to be flushed. And also the context switching itself, of course.
However, you can still do a lot on Linux to equalize things if you really want to get down to it. For anything but special cases Linux really does a good job with scheduling. After all, you are likely not running much else other than your intended service.
That said, for me this thread was a slight wakeup-call that made me look more into fibers and co-routines. I have been wanting to use these for a long time for some things.
1M cycles is roughly 300 microseconds (assume 3 GHz processor, so 3 cycles is 1 nanosecond). Eli’s graph from the post I referenced above, has a context switch in the 1-3 microsecond range [1] depending on taskset/core pinning. The high end (3 microseconds) is about 10000 cycles then.
Maybe you mean fork() or pthread_create for your 1M cycles?
[1] https://eli.thegreenplace.net/images/2018/plot-launch-switch...
Some types of software optimizations require the ability to correctly infer local CPU cache contents, which is difficult when arbitrary processes are semi-randomly stepping all over that cache.
Like only scheduling work on logical cores that share a physical core after all physical cores have a busy logical core (I.e. fill up the even cores first).
To your point, it is a double-edged sword. Writing your own schedulers requires a much higher degree of sophistication than using the one in the OS. It is a skill that takes a long time to develop and requires a lot of first principles thinking, there is loads of subtlety, you can't just copy something you found on a blog. It also isn't just about being able to predict the behavior of your workload better than the OS, you can also adapt your workload to the schedule state since it is exposed to your application, the latter being a greatly overlooked capability.
Once you know how to design software this way, it not only generates large increases in throughput but also enables many elegant solutions to difficult software design problems that simply aren't possible any other way. While the learning curve is steep, once you are accustomed to writing software this way it becomes pretty mechanical.
Exactly. And some apps automatically do that by default, as if they are so arrogant as to think they must be the only program running on that machine. Maybe good for a server app, terrible advice for a general desktop app.
C# async/await is pretty good for this (IO operations do not count towards CPU task count).
Even if you pre-create a thread (thread pool), when the task is small enough (less than 1,000 cycles), it is less expensive to do it in place (for example, with fibers), because of the cost of context switching.
Userspace threads are more light-weight, but probably still worse than just using fibers and co-routines. Depends on your needs, I suppose.
A fiber switch can be done in less than 10 clock cycles.
It seems at least the Linux folks optimized the crap out of clone() over the last years.
The most essential Linux benchmark is compiling the Linux kernel (since it's something the Linux kernel developers do all the time, so they really feel the impact). The clone() system call is used both to create new threads and to create new processes, and the Linux kernel compilation uses a large amount of short-lived processes (each C file is a new C compiler process). It's only natural that clone() is heavily optimized, together with the filesystem caches (each new C compiler process reads the source code files from scratch).
When I ran benchmarks to compare a thread-per-client model to a single-threaded, event-based one, the single-threaded throughput was around 2 to 3 times higher for as few as 1000 clients.
$taskset --cpu-list 8 ./costofthread avg: 11000~
$taskset --cpu-list 8,11 ./costofthread avg: 33000~
$./costofthread avg: 60000~
What programming languages' de-facto thread implementations are not wrappers around pthreads? I think Go has its own thread implementation? Or am I mistaken?
> A newly spawned Erlang process uses 309 words of memory in the non-SMP emulator without HiPE support. (SMP support and HiPE support both add to this size.)
And a word is the native register size, so 4 or 8 bytes these days, so fairly small, but not 64 bytes small.
- https://wiki.haskell.org/Parallelism#Multicore_GHC
The upcoming Project Loom, intends to make it so that green threads become the default (aka virtual threads on Loom), but you can still ask for kernel threads, given that is what most JVM implementations have converged into.
With all possible the learnings from Java, .NET, Erlang, TBB, Concurrency Runtime, and yet ISO C++ did not manage to get a proper concurrency story, and it full of traps like the one you mention.
Another one is std::async, which might actually be synchronous, depending on a set of factors.
I’ll also be interested to see the same benchmark but using pthread_create directly.
An equivalent to TBB or GCD will be in C++23 std libraries, but you can often do better with coroutines, in 20.
TBB and GCD still need to sychronize sometimes, and they randomize workload assignment, which is bad for cache locality (i.e. bad). If you can arrange static assignment and avoid need to synchronize, you can do better, sometimes much better.
See, for instance, https://www.theimpossiblecode.com/blog/intel-tbb-on-raspberr....
Other than filesystem, back in a 2010 project, I have never used Boost.
I have enough concurrency in C++, thanks Concurrency Runtime.
Somehow it is ironic to see C++ catching up with java.util.concurrency and TPL.
Java TPL is very, very limited when compared to C++23 executors.
"Concurrency", by the way, has come to refer to the synchronizing interactions that cause slower-than-xN parallelism.
Also other programming languages also don't stand still.
The mode chosen by the C++ committee is to favor development of library features outside the Standard, and then adopt the successes. That entails waiting to see what is a success, and relying on non-Standard library implementations in the meantime. The '23 executors design has been a long time coming, but is overwhelmingly better -- meaning, applicable to a much broader space of execution models -- than early designs, without compromise on performance. Other languages routinely compromise on performance, which is often the right choice for them.
* Do your tasks block? How many threads do you need to make sure you can use all your CPUs.
* Do your tasks access different sets of memory? Would keeping similar tasks on the same CPUs reduce cache misses.
* Do your tasks have different priorities? You might need a pool for each priority.
For a UI program that isn’t doing anything really intensive or real-time, having a common thread pool makes a lot of sense, and can reduce resource use (stacks add up once you get to many 10s or 100s of threads...), and improve latency (a work queue with many threads will get more CPU than another with the same amount of work but fewer threads)
I used nodejs for a project, and assumed that "it's all javascript on one thread" would leave threading issues behind.
My application curiously stopped responding whenever I had 5 or more users. Connected users could continue to do anything, but new users couldn't connect, and existing users sessions would hang when executing any code that wrote to a logfile, making debugging even harder. Using the nodejs debugger, the internals of write(...., cb) were just never calling the done callback.
After hours of head scratching I found that most IO from nodejs is not asynchronous and callback based as the docs suggest, but is in fact blocking IO done from worker threads. My process was using pipes to communicate with other processes, and those pipes were doing blocking writes, and when blocked, the worker thread was blocked.
There are 4 worker threads by default, so whenever 5 users were using the system, all worker threads were tied up and it would fail. It would have been nice for nodejs to at least have printed to the console "All worker threads busy for >1000ms. See nodejs.com/troubleshooting/blockingfileio.htm" or something.
* Do you have sufficiently large batches that you can efficiently assign to one thread?
If not, then you're just wasting a lot of time waking up to receive inputs, assigning them to threads (-> put them on a work queue or similar, with all the locking / atomics), and waking up a thread to pull an item (locking / atomics), process it, go to sleep...
It's easy to end up spending more time juggling tasks and switching tasks than performing any useful work.
I personally prefer to do as much as possible just in one thread, where you can run things asynchronously with a single threaded message loop and then have a thread pool next to that for expensive computations. This also tends to reduce the number of things that need to be protected with a mutex.