I designed and implemented a dynamic scheduler for a streaming dataflow language a few years ago [1]. We wanted a runtime system which could have hundreds of OS-level threads execute thousands of dataflow operators that communicate in a dataflow manner. Threads should not be statically assigned portions of the dataflow graph so that we could elastically add or remove OS-level threads based on observed performance.
We compared that scheduler to some other options, including just giving every operator in the graph its own dedicated thread. One test application was a simple 1,000 operator pipeline. We used two machines, one with 176 cores, the other 184 cores. To my surprise, with the pipeline application, the dedicated thread model beat my fancy scheduler in raw performance by up to a factor of 2. Keep in mind that that's 1,000 threads, all doing work, on machines with only 176 and 184 cores.
Of course, you would not want to do this in practice, even though the raw performance was so high: the machine was so massively oversubscribed during such experiments that it could barely keep up with a simple interactive shell.
But, my intuition had been wrong: I had thought that surely having 10x the number of threads as cores would mean the overall performance would crawl because of context switching time. It did not. See section 5.1 of my paper below for the experiment.
[1] Low-Synchronization, Mostly Lock-Free, Elastic Scheduling for Streaming Runtimes, PLDI 2017, https://www.scott-a-s.com/files/pldi2017_lf_elastic_scheduli...