The easiest way to understand 100,000 concurrent threads is by making all 100,000 threads do the same thing and execute the same code.
EDIT: Works on the small scale too. If you're using a CPU, then you can assign exactly 128-threads to be your data-structure's "workers". You access the data-structure sequentially, but all manipulations are done in parallel by the 128x "worker" threads in the background.
Assuming a 64-core / 128-thread CPU like an AMD EPYC or whatever.
Normally, your data-structure "commands / frontend" aren't the bottleneck, instead is the data-structure's "work" that bottlenecks. As long as that work happens in parallel, you're probably fine.
------
Bonus points: this leads to "obvious" NUMA-locality. If your NUMA-cluster is just 8c/16-threads, you can have a NUMA-local set of 16-workers with affinity set. You scale to the level of your L3 cache (or whatever arbitrary NUMA-node you want), and different threads can have their "local workers" to choose from.