All about thread-local storage
maskray.me
maskray.me
I had known that thread-local variables can be access pretty fast, via dedicated segment register, but I was not clear how can one make this work for dynamically loaded PIC code, like most .so files.
Turns out you can't. You only get fast access via dedicated register if you are using variable declared in the main program. The .so files have to call special function which does multiple memory lookups to get the actual location, probably severely reducing performance.
(and this is another case when seemingly simple operation -- getting variable value -- gets internally translated to dozens of operations and a function call)
I just had this in mind because skimmed it a few hours ago in the context of https://ssrg-vt.github.io/hermitux
mentioned in https://news.ycombinator.com/item?id=26142285
Is Parallel Programming Hard, And, If So, What Can You Do About It?
https://cdn.kernel.org/pub/linux/kernel/people/paulmck/perfb...
if I wanted to store a per-thread count in rust, it would make zero sense to use TLS for that, it would just be on the stack in the context of the thread's lexical scope:
let mut threads = Vec::new();
let global_count = Arc::new(AtomicUsize::new(0));
for _ in 0..n_threads {
threads.push(std::thread::spawn({
let global_count = global_count.clone();
move || {
let mut thread_count = 0;
for _ in some_iteration {
thread_count += 1;
}
global_count.fetch_add(thread_count, Ordering::Release);
}
});
}
for handle in threads { let _ = handle.join().unwrap(); }
println!("global count = {}", global_count.load(Ordering::Acquire));
that is a form of "thread-local storage", I guess, but does not involve any of the TLS primitives.For high quality code you often pass around the RNG instance explicitly (improves testability), but often I just want a random number without bothering with explicit state management.