The way we do TLS now is too fiddly and accommodating of legacy code. What to keep TLS simple and fast? Treat it like an arena.
1. On each thread startup, including the main thread, carve out a huge chunk of address space, say 1GB, for that thread's TLS arena
2. on dlopen (or main program startup), allocate each loaded DSO's TLS out of the arena. Fail in the unlikely case that we run out of address space
3. on dlclose, recover committed memory using MADV_FREE on any now-unused parts of the TLS arena (once for each thread
4. to access a thread local, pull the offset of the variable out of the GOT and offset into the per thread arena. Nice and simple.
Does this approach waste address space? You bet. Will it work on 32 but systems? Absolutely not. Is it simple, fast, and robust? Yes.