0+0 > 0: C++ thread-local storage performance
yosefk.com
yosefk.com
Separately, the post describes one of a few aspects of optimizing funtrace (https://github.com/yosefk/funtrace), which I think is the fastest open-source function tracing profiler for C++ today, and which can be ported relatively easily to Rust/other languages producing native code (the runtime is ~1200 LOC of C++; the trace decoder would need very few changes.) A description of how funtrace works as well as why you'd want a tracing profiler (as opposed to a sampling profiler like perf) and how you'd use it is here: https://yosefk.com/blog/profiling-in-production-with-functio...
I think what you want is to force the "initial exec" model using an attribute to get the more efficient code.
IIRC this (or something equivalent) is what libGL does because basically every OpenGL function call needs to read the thread-local variable holding the current GL context.
The downside is that dlopen()ing your library may fail.
If initial-exec TLS does not work due to the dlopen issue, on x86-64 and recent-enough distributions, you can use -mtls-dialect=gnu2 to get a faster variant of __tls_get_addr that requires less register spilling. Unfortunately glibc and GCC originally did not agree on the ABI: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=113874 https://sourceware.org/bugzilla/show_bug.cgi?id=31372 This has been fixed for RHEL 10 (which switched to -mtls-dialect=gnu2 for x86-64 for the whole distribution, thereby exposing the ABI bug during development). As the ABI was fixed on the glibc side in dynamically-linked code, the change is backportable, but it's a bit involved because the first XSAVE-using change upstream was buggy, if I recall correctly. But the backport is definitely something you could request from your distribution.
Note that there was a previous bug in __tls_get_addr (on all architectures that use it), where the fast path was not always used after dlopen: https://sourceware.org/bugzilla/show_bug.cgi?id=19924 This bug introduced way more overhead that just saving registers. I expect that quite a few distributions have backported the fix. This breaks certain interposed mallocs due to a malloc/TLS cyclic dependency, but there is a workaround for that: https://sourceware.org/git/?p=glibc.git;a=commitdiff;h=018f0...
The other issue is just that the C++ TLS-with-constructors design isn't that great. You can work around this in the application by using a plain pointer for TLS access, which starts out as NULL and is initialized after a null check. To free the pointer on thread exit, you can use a separate TLS variable or POSIX thread-specific data (pthread_key_create) to register a destructor, and that will only be accessed on initialized and thread exit.
This sort of question is probably more suited to libc-help: https://sourceware.org/mailman/listinfo/libc-help/
This feels like the perfect situation to preallocate a gigabyte or something of virtual memory for extending the TLS, similar to how the stack is. But, testing on my system, looks like the allowed initial-exec TLS size is just ~1700 bytes.
Frankly, if your name isn't `libGL.so` you shouldn't even try to mix initial-exec with dlopen. Just link your libraries normally dammit!
dlopen is a requirement for importing native libraries in non-compiled languages; and, regardless, I as a library author don't get to choose whether users will avoid using dlopen and so have to assume worst-case.
I don't know how the kernel manages it internally, but there's no need for PROT_NONE preallocated virtual memory to be mapped to actual CPU-accessible pages at least; and `mmap(NULL, 1ULL<<46, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0)` takes ~4 microseconds to map 64 terabytes of virtual memory so it's definitely not 0.002x overhead. (perhaps the overhead amount changes depending on how close to a page level the size is, but it shouldn't be too much regardless)
This'd essentially be turning the preallocated TLS space as a memory allocation arena (and you could actually even just choose to provide an alloc+free interface for programs to dynamically allocate fs-relative-offsets to use for custom threadlocals?).
(then there's general problematicness of virtual memory; such PROT_NONE never-touched memory still counts towards virtual memory usage, which is annoying; browsers/Java/etc already suffer from this, but it'd be rather ugly for literally all processes to have such. I'd quite like a memory usage counter that includes all memory that is or ever was writable, but not PROT_NONE never-touched; i.e. how much memory the process can eventually require without running explicitly requesting more via syscalls, but afaik such just doesn't exist, or at least isn't a standard-displayed thing)
This concept is called "commit charge". Windows MM models it explicitly. Linux ought to as well. I agree it's a more useful concept than just address space allocated.
MADV_FREE doesn't affect anything afaict.
An alternative that might be worth looking into is just hashing the FS/GS into a table index. It will be slower than the well-optimized case, but it will let you opt out of the TLS allocation process altogether. This might be a good thing in some cases for a low-level facility like a function tracer.
Note that my question is about shared libraries. If the thread_local is linked into an executable, I guess you could indeed save the offset somewhere and then add the value of %fs to it, though if this is a way to work around the constructor issue, I prefer to not have a constructor. The question is if this sort of direction can help for thread-local storage allocated by a shared library.
Allocate the variable normally, then compute the offset in one thread (e.g. offset = uintptr_t(&variable) - get_fs()), then access it by adding the offset to FS in any thread (e.g. (vartype *) (offset + get_fs())). The only difference from how it normally works is that you can manually force it to be inlined, sidestepping the codegen problems you described in your post. But if you can avoid those problems by not using constructors instead, that's definitely better.
I used "FS/GS" because GS is used instead of FS on some systems for the same purpose.
The shared library-specific issues are one of the reasons I was suggesting maybe looking into hashing, e.g. perhaps as a fallback solution when the TLS approach fails.
As a result of the later kernel support, it may not be for everyone to turn on. Until then, the replacement on GNU Linux is just to load %fs:0, which the x86-64 ABI requires to have the same value (or %gs:0 on i386). However, usually, for initial-exec TLS access, the address of variable is not required to be in a register.
Is SEGFS/%fs-based access slower than loading the base address with RDFSBASE (which may require spilling a register) and then using base+offset access? I haven't seen such reports.
- LTTng tracer
- tcmalloc
Curious if there are other prominent users of rseq.
Funtrace on the other hand does support ftrace for tracing context switches (https://yosefk.com/blog/profiling-in-production-with-functio...), but it doesn't require ftrace for tracing function calls made by your threads (the problem with ftrace as well LTTng's kernel modules being, of course, permissions; which shouldn't be a problem in any reasonable situation by my standard of "reasonable", but many find themselves in unreasonable situations permissions-wise, sadly.) So I don't think funtrace can use rseq, though I might be missing something.
In the process of investigating this, I also realized that there's a ton of other unique-per-thread pointers accessible from that structure, most notably including the value of %fs itself (which is unfortunately unobservable afaict), the address of the TCB or TLS structures, the stack guard value, etc. Since the goal is just to have a quickly-readable unique-per-thread value, any of those should work.
Windows looks similar, but I haven't investigated as deeply.
[0] https://github.com/andikleen/glibc/blob/b0399147730d478ae451...
[1] https://github.com/andikleen/glibc/blob/b0399147730d478ae451...
How do you avoid the problem of threads migrating between CPUs at arbitrary instruction boundaries?
rseq is short for restartable sequences; you mark beginning and end of a range of instructions, plus an abort path. The kernel checks this during task-switch and if you're within the range the instruction pointer is changed to the abort path on resumption.
(It's called restartable because the assumption is that the abort path will try again. Or at least, aborting the block of instructions midway through is recoverable.)
The primary limitation is that the "result" of the block in most cases needs to be concentrated into the last instruction of the block (similar to an atomic release write.) Otherwise you'd need to somehow recover from partially executed rseq blocks having changed some visible state but not fully completed.
Question about your build system comment: my build system doesn't stoop to figuring out if `-fPIC` is needed, but it also doesn't add it unless the user asks. Were you talking about that or build systems that add it automatically?
I don't think build systems add -fPIC automatically, nor remove it automatically. C++ build systems do not stoop to the question of how to best build a C++ program, by and large. They are more task graph executors - either bad ones, like make, or good ones, like Bazel, but mostly task graph executors; the most "C++ support" you will get is native support for scanning #include files (as opposed to doing it yourself like make forces you to.)
I want to build a "standard library" for my build system that would stoop to that. I can do that because my build system is not just a task graph executor; it is backed by a full programming language and can add its own libraries to that language. IOW, I can add a `cpp` package to the build system that implements support for how to best build a C++ program.
So if you have a wishlist for that support, I'd love to hear it.
GNU Autotools do, if you tell, that you have a shared library. You can also switch between static/dynamic linking at build time and if you also use GNU Libtool, it will figure out the flags of your build platform at build time.
> as opposed to doing it yourself like make forces you to
You can definitely don't have to do that. GNU Automake does it by default, but if you are using plain Make, you can also use makedepend or the appropriate flags of your compiler.
> either bad ones, like make
What is wrong/missing with make as a task graph executor?
> or good ones, like Bazel
What can Bazel do better?
You can also use the PTWRITE instruction to attach metadata to the stream which seems very powerful.
Hope we can see such an extension on AMD as well.
Typically you get a cycle count every six branches, give or take.
https://yosefk.com/blog/profiling-in-production-with-functio...
https://danluu.com/perf-tracing/
Regarding the slowdown - magic-trace reports 2-10% slowdowns which IMO is actually fine even for production (unless this adds up to a huge dollar cost, for most people it won't) since in return for this you are actually capable to debug the rare slowdowns which are the worst part of your user experience.
However, the hardware feature that I propose (https://yosefk.com/blog/profiling-in-production-with-functio...) would likely have lower overhead since it relies on software issuing tracing instructions, eg at each function entry & exit (rather than any control flow change), and it could be variously selective (eg exclude short functions without loops; and/or you could configure the hardware to ignore short calls. BTW maybe you can with Intel Performance Trace, too, I'm just not really familiar with it.)
Like I said there, I'm frankly shocked that all CPUs haven't raced to implement similar features, that magic-trace which is built on top of Intel Performance Trace isn't used more widely, and that developers aren't insisting on running under magic-trace in production and requiring to deploy on Intel servers for that purpose.
The extension I propose is much simpler, and seems similar to what PTWRITE would do if it was the only feature in Intel Performance Trace. I have a lot of experience in chip architecture, and I believe that every CPU maker and every chip maker can support this easily - much more so than full feature parity with Intel Performance Trace. I hope they will!
I wonder if this is a general issue relating to memory ordering or out-of-order execution, or whether this can be implemented more efficiently in a different extension.
Thank you for the linked article! Agreed on the huge potential for using these tools in production. The community could definitely benefit (even indirectly) by pushing for this kind of instruction set more widely.
> [...]
> I don’t know how to generalize the principle to make it explicit and easy to follow.
Coming from mathematics that is what I would call using the right level of abstraction.
If you want to prove 0 + 0 = 0 and you're getting tied up with stuff like how the direct sum of two Cauchy sequences should converge to the sum of the two limits then you're not working in the right level of abstraction. You're not supposed to know about Cauchy sequences yet if all you're given is the neutral element for addition.
In some rare cases it can help knowing about sublevels of abstraction. Such as the difference between a general linear space and one equipped with an inner product. Just because you can make an inner product doesn't mean you should, and if you don't you'll find some arguments a lot easier because you're not distracted by stuff like adjoints and orthonormal basis vectors etc. (one side effect is that gradient descent no longer works, and you really ought to know why). You can do similar things by refusing to decide on a coordinate system.
1. On each thread startup, including the main thread, carve out a huge chunk of address space, say 1GB, for that thread's TLS arena
2. on dlopen (or main program startup), allocate each loaded DSO's TLS out of the arena. Fail in the unlikely case that we run out of address space
3. on dlclose, recover committed memory using MADV_FREE on any now-unused parts of the TLS arena (once for each thread
4. to access a thread local, pull the offset of the variable out of the GOT and offset into the per thread arena. Nice and simple.
Does this approach waste address space? You bet. Will it work on 32 but systems? Absolutely not. Is it simple, fast, and robust? Yes.
> Block variables with static or thread(since C++11) storage duration are initialized the first time control passes through their declaration (unless their initialization is zero- or constant-initialization, which can be performed before the block is first entered). On all further calls, the declaration is skipped. --- https://en.cppreference.com/w/cpp/language/storage_duration
From https://maskray.me/blog/2021-02-14-all-about-thread-local-st...
> If you know x does not need dynamic initialization, C++20 constinit can make it as efficient as the plain old `__thread`. [[clang::require_constant_initialization]] can be used with older language standards.
Regarding `data16 lea tls_obj(%rip),%rdi` in the general-dynamic TLS model, yeah it's for linker optimization. The local-dynamic TLS model doesn't have data16 or rex prefixes.
Regarding "Why don’t we just use the same code as before — the movl instruction — with the dynamic linker substituting the right value for tls_obj@tpoff?"
Because -fpic/-fPIC was designed to support dlopen. The desired efficient GOTTPOFF code sequence is only feasible when the shared object is available at program start, in which case you can guarantee that "you would need the TLS areas of all the shared libraries to be allocated contiguously:"
# x86-64
movq ref@GOTTPOFF(%rip), %rax
movl %fs:(%rax), %eax
With dlopen, the dynamic loader needs a different place for the TLS blocks of newly loaded shared libraries, which unfortunately requires one more indirection.Regarding "... and I don’t say a word about GL_TLS_GENERATION_OFFSET, for example, and I could."
`GL_TLS_GENERATION_OFFSET` in glibc is for the lazy TLS allocation scheme. I don't want to spend my valuable time on its implementation... It is almost infeasible to fix on the glibc side.
Thanks - I didn’t realize this was mandated by the standard as opposed to “permitted” as one possibility (similarly to how eg a constructor of a global variable can be called before main or upon first use or anywhere in-between according to the standard). Updated the post with this point
> The desired efficient GOTTPOFF code sequence is only feasible when the shared object is available at program start, in which case you can guarantee that “you would need the TLS areas of all the shared libraries to be allocated contiguously”
Indeed I didn’t mention -ftls-model=initial-exec originally (I now added it based on reader feedback; it can work when it will work, which for my use case is a toss-up I guess…), but my point is that you could allocate the TLSes contiguously even if dlopen was used, and I describe how you could do it in the post, albeit in a somewhat hand-wavy way. This is totally not how things were done and I presume one reason is that you don’t carve out chunks of the address space for a use case like this as described in my approach - I just think it would be nice if things worked that way.
3.7.2/2 [basic.stc.thread]: A variable with thread storage duration shall be initialized before its first odr-use (3.2) and, if constructed, shall be destroyed on thread exit.
This allows the constructor to be called at any point before the first use, similarly to "normal" globals, though implementations made different tradeoffs in these 2 cases
Edit: ah, see https://www.akkadia.org/drepper/tls.pdf which I was clued into by https://news.ycombinator.com/item?id=43079061
The situation is much more complicated than it sounds, because the question of where the variables are isn't known at compile time. It has to be done at link time. Which may include dynamic link time as a library is loaded. That all sounds fairly horrible.
All to handle static constructors. If you don't have thread-static variables, or they don't have constructors, it's much simpler.
Dynamic linking is a whole other beast, and maybe for thread_local's in a DLL it's acceptable (or the only possible way?) to do construct them lazily - if you're using DLL+TLS+global constructors you deserve the pain.
Also creating a thread, doing little work and exiting immediately is a bad use of threads -- I don't think that everyone using thread_local should be punished to better support poorly written programs. Most threads are long-running and most programs spin up a bunch of worker threads at initialization time, I'd argue variable access time is way more important than thread startup time.
Do agree that threadlocal read speed should be a lot more important, but it'd quite suck if the "acceptable performance" thread usage approaches would vary dramatically depending on what libraries you've (or something else) has unrelatedly loaded. (though maybe other overheads might appear from other sources similarly, keeping this not the most important slowdown, idk)
This is not true, at least on Linux. With appropriate settings (e.g. a small stack) thread creation can be extremely cheap.
As far as why it's worse with fPIC/shared: there are a variety of TLS models: General Dynamic, Local Dynamic, Initial Exec, and Local Exec. And they have different constraints / generality. The more general models are slower. IIRC IE/LE won't work with shared library thread_locals, but it's been a while so don't quote me on that.
Generally agree that it seems like the compiler could in theory be doing a better job in some of these circumstances.
> I’m sure there’s some dirty trick or other, based on knowing the guts of libc and other such, which, while dirty, is going to work for a long time, and where you can reasonably safely detect if it stopped working and upgrade it for whatever changes the guts of libc will have undergone. If you have an idea, please share it!
Find the existing TLS allocations, hope there's spare space at the end of the last page, and just map your variables there using %fs-relative accesses?
Always fun to see another high performance tracing implementation. We do some similar things at work (thread-local ringbuffers), though we aren't doing function entry/leave tracing.
> in my microbenchmark I get <10 ns per instrumented call or return
A significant portion of that is rdtsc time, right? Like 50-80%. Writing a few bytes to a ringbuffer prefetched in local cache is very very cheap but rdtsc takes ~4-8 nanos.
> While we're on the subject of snapshots - you can get trace data from a core dump by loading funtrace_gdb.py from gdb
Nice. We trace into a shared memory segment and then map it at start time, emitting a pre-crash trace during (next) program start. Maybe makes more sense for our use case (a specific server that is auto-restarted by orchestration) than a more general tracing system.
Exactly right, though it seems totally insane to me coming from an embedded background where you get the cycle count in less than 1 ns but writing to the buffer would be "the" performance problem (and then you could eg avoid writing short calls as measured at runtime, but on x86 you will have spent too much time on the rdtsc for this to lower the overhead.) There's also RDPMC but it's not much faster and you need permissions(tm) to use it, plus it stops counting on various occasions which I never fully understood.
Regarding prefetching - what do you do perfetching-wise that helps performance?.. All my attempts to do better than the simplest store instructions did nothing to improve performance (I tried prefetchw/__builtin_prefetch, movntq/_mm_stream_pi and vmovntdq/_mm_stream_si128, all of them either didn't help or made things even slower)
Absolutely nothing -- the CPU internally just does a very good job predicting the ringbuffer write pattern, for obvious reasons.
Actually, thread globals only being used in some threads is probably the common case. Remember that threads are not only created by the main executable but often also by libraries. Keeping per-thread overhead low until the TLS variables are actually needed sounds like a good design goal.
> funtrace sidesteps the TLS constructor problem by interposing pthread_create, and initializing its thread_locals in its pthread_create wrapper
Sounds extremely fragile. What if someone calls clone() directly or another (possibly new) function to create a thread.
In any case, one point of funtrace is that it's small (~1K LOC runtime) and you can tweak it easily, including calling its TLS init code from your threads if you don't create them with pthread_create "like most people" - even without this issue, people do "green threads" with setcontext/getcontext and who knows what else, which will need its own hacks to support, and so I think that an easily hackable runtime is a good thing given how hard it is to do this in a one-size-fits-all fashion.
As a counterexample, LLVM XRay, a tracing profiler from Google which is at least 10x bigger than funtrace (if you don't count the compiler instrumentation they introduced), took almost a decade to gain shared library support, and it's not done yet. So I think "small and hackable" has its advantages.
I'm not sure how many people interested in this article is interested in C++ compile times but I once measured and wrote an article https://bolinlang.com/wheres-my-compile-time
I'm also mostly familiar with Windows, and on Windows until recently (few years ago), dynamically loading (e.g. with "LoadLibrary", e.g. "dlopen") a .dll caused issue with the .dll's own thread_locals. Microsoft fixed this, but folks have observed slower code
https://developercommunity.visualstudio.com/t/5x-performance...
To quote only the observed case there:
``` In VS2017 15.9.26 this executes in ~270ms and with VS2019 16.7.1 it takes ~1450ms. ```
Here are the notes too - https://learn.microsoft.com/en-us/cpp/overview/cpp-conforman...
Do you know what these issues were? I'm curious because I'm working on Pd (https://en.wikipedia.org/wiki/Pure_Data), which uses lots of thread local variables internally when built as a multi-instance library. Libpd itself may be loaded dynamically when embedded in an audio plugin. I'm not aware of any problems so far...
But then Microsoft added /Zc:tlsGuards - https://learn.microsoft.com/en-us/cpp/build/reference/zc-tls... - which is now the default that fixes the issue, but with some significant performance penalty (e.g. the "bug" that I've listed).
I guess you can't have it both ways easy...
On the clang/clang-cl side, there is https://clang.llvm.org/docs/ClangCommandLineReference.html#c...
to support this.
So check your compiler version and options :)
Also the notes posted here about CRT mixing might apply to you (not sure though) - https://learn.microsoft.com/en-us/cpp/porting/binary-compat-...
I work in a gamedev world, and plugins, ffi, delay loaded dlls etc. are constant pain that one needs to look and solve issues around.
Do you happen to have a link to the original MSVC bug report (i.e. the wrong thread locals, not the performance regression)?
https://learn.microsoft.com/en-us/cpp/overview/cpp-conforman...
and also the implementation in "clang" (for "clang-cl" being conformant with MSVC) - https://reviews.llvm.org/D115456#3217595
then last year clang-cl also added ways to disable this (if need to), probably this hit some internal issue and had to be resolved. Maybe "thread_local" have become more widely used (unlike OS specific "TlsAlloc")
I made an elegant program that counts from zero to zero. It took me a week.
https://www.akkadia.org/drepper/tls.pdf is the go-to reference for TLS models, cf. first paragraph of each subsection of section 4.
-fvisiblity=hidden / __attribute__ ((visibility ("hidden"))) doesn't seem to do anything.
Are you sure that either of these speeds up TLS access for thread_locals defined in shared libraries?
Also note that these sequences are highly architecture dependent and cost as well as cost differences will vary e.g. for ARM or ARM64.
I stick to the broad conclusion that thread_locals without constructors linked into an executable rather than a shared library are the fastest and most performance-portable by far, but the visibility point is very worth mentioning.
7e2f0: f3 0f 1e fa endbr64
7e2f4: 64 48 8b 04 25 f8 ff mov %fs:0xfffffffffffffff8,%rax
7e2fb: ff ff
7e2fd: c3 ret
7e2fe: 66 90 xchg %ax,%ax
But they did not inline the function that does that, even with `__attribute__((always_inline))` and link-time optimization. I need to investigate that.I am grateful that I only use one multi-call executable (built from a monorepo made partially because of your post, btw) because I think shared libraries in C might have the same sort of problems, just without constructors.
Tip to the author: you can implement a lock-free buffer pool cache for as many threads you need in ~200 LoC. Another tip: it's not gonna be faster but you may get more control over the resources
What I see here in the article is pure theoreticizing and dwelling over details which in 99.9% of the cases do _not_ matter. There wasn't a single experiment shown by the author by which he shows that his use-case is in that 0.01%.
I was basically done with reading the article when I read the following nonsense
> accessing an extern thread_local with a constructor involves a function call, with the function testing the guard variable of the translation unit where the thread_local is defined.
And then
> But with inlining, the fast path is quite fast on a good processor.
Right, the fast path of inlined ctr call vs slow path of non-inlined ctr call. Really? I call this a BS, especially without the experiments showing the data.
The unfortunate effect of this and similar blogs is that tomorrow you will have to deal with someone at your work who happened to read this and such material on the web and will take it for granted. He will take it further and implement an "optimization" yet when you ask to demonstrate the problem he was solving you will get big nothing. Just as in this case.