>The solution is for only one of the two layers to parallelise
If you have a common scheduling API, you can manage this much more elegantly. For example make(1) can control the concurrency level across recursive invocations by using a "job server" https://www.gnu.org/software/make/manual/html_node/Job-Slots...
With this, you can have, for example, 3 top level subprocesses, each spawning multiple threads of their own, but never exceeding the CPU count.
Alternatively, parallel could make it's subprocesses think that they're running on a 1-core machine, although this may have some subtle side affects.