The easiest way to understand 100,000 concurrent threads is by making all 100,000 threads do the same thing and execute the same code.
EDIT: Works on the small scale too. If you're using a CPU, then you can assign exactly 128-threads to be your data-structure's "workers". You access the data-structure sequentially, but all manipulations are done in parallel by the 128x "worker" threads in the background.
Assuming a 64-core / 128-thread CPU like an AMD EPYC or whatever.
Normally, your data-structure "commands / frontend" aren't the bottleneck, instead is the data-structure's "work" that bottlenecks. As long as that work happens in parallel, you're probably fine.
------
Bonus points: this leads to "obvious" NUMA-locality. If your NUMA-cluster is just 8c/16-threads, you can have a NUMA-local set of 16-workers with affinity set. You scale to the level of your L3 cache (or whatever arbitrary NUMA-node you want), and different threads can have their "local workers" to choose from.
Maybe, but it surely is inconvenient!
When our program can use pipe-line with several pipelines for concurrency (like GPU shaders work), sure, that's a good way to go. However the majority of what threads are doing is not pipe-line friendly work.
When we want our program to do different things at the same time (which is what we mostly use threads for), trying to do it in a pipe-line-like manner won't work.
For example, serving multiple http requests at the same time - some will be a POST to an API endpoint, some will be a GET for a static file, some controller code will make multiple requests to different backend microservices (each request in a different thread), etc.
Or maybe some client code, that makes several concurrent requests for different resources to different destinations, while each thread still needs to update the UI (progress meter, for example).
Or a parser for input from some upstream processes: you want to parse in parallel if at all possible.
Or the program in a microcontroller that cannot miss bytes coming in on one bus while it is sending bytes out on another bus.