Yeah, a gigabyte is most likely extremely overkill indeed, a megabyte or so would be plenty; though the goal would be to get threadlocals to be able to be as arbitrarily large as non-initial-exec threadlocals so it wouldn't break anything ever.
I don't know how the kernel manages it internally, but there's no need for PROT_NONE preallocated virtual memory to be mapped to actual CPU-accessible pages at least; and `mmap(NULL, 1ULL<<46, PROT_NONE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0)` takes ~4 microseconds to map 64 terabytes of virtual memory so it's definitely not 0.002x overhead. (perhaps the overhead amount changes depending on how close to a page level the size is, but it shouldn't be too much regardless)
This'd essentially be turning the preallocated TLS space as a memory allocation arena (and you could actually even just choose to provide an alloc+free interface for programs to dynamically allocate fs-relative-offsets to use for custom threadlocals?).
(then there's general problematicness of virtual memory; such PROT_NONE never-touched memory still counts towards virtual memory usage, which is annoying; browsers/Java/etc already suffer from this, but it'd be rather ugly for literally all processes to have such. I'd quite like a memory usage counter that includes all memory that is or ever was writable, but not PROT_NONE never-touched; i.e. how much memory the process can eventually require without running explicitly requesting more via syscalls, but afaik such just doesn't exist, or at least isn't a standard-displayed thing)