If your threads don't block and you have one main application running on the node, you usually just want to run one thread per node and you're done. If you are running a very network heavy load where you can eliminate or highly reduce cross thread communication, you may want one core per NIC tx/rx queue and one thread per core to eliminate cross-core communication; any cores above the number of queues will just be idle, because cross-core communication is more expensive than the work they can do (but that's not a super common scenario).
A control system to add and remove nodes makes sense if you're cloudy, though, since there's a cost for running nodes.
That would be my intuition too, in particular that usually the cost of idle workers is pretty low, so it's better just to preallocate some fixed max number of workers than try to scale them.
I wouldn't necessarily intuit the cost of idle workers is low, more that the non-cpu cost of workers is roughly fixed, and if it's too expensive to run more than you need at idle, it's still going to be too expensive to run that many at full load. Sometimes it's hard to know what the max load capacity is, but Apache configs where the worker count scales in and out are really easy to get into load is high -> spawn more workers -> use too much memory -> pick your poison: evict too much disk cache / swap to death / oom killer
[0]: https://people.eecs.berkeley.edu/~brewer/papers/SEDA-sosp.pd...
It might be more suited to infinitely-scalable situations like cloud VMs, where you are trying to optimize money spent.
Right, it should base the feedback on throughput measures instead. If throughput starts dropping, then the threadpool is overcommitted, and should scale back.
I believe the SEDA papers mentioned in the README discusses this as backpressure.