Maybe not any better syntax, but xargs is usually installed by default in most distros, and I've never found parallel to be advantageous over it. I don't think this particular wheel needs to be invented again.
... | xargs -P10 -i sh -c 'command {} 1>{}.out 2>{}.err'
or maybe with tee:
... | xargs -P10 -i sh -c 'command {} | tee -a {}.out'
It is definitely not something I would dare to run in production.
xargs is part of POSIX, so it should be everywhere.
(-P/-n are not, but both GNU and BSD versions do support it.)
I only have quad core, but -P values higher than 4 give better performance. Maybe because the tests take different amounts of time, and it avoids them backing up. To get x7, I had -P equal to the number of tests.
real 0m1.818s
user 0m2.020s
sys 0m0.900s
top. (It completes in 1 sec, so put in a while loop - BTW top isn't behaving normally, "1" not giving multiple CPUs and not fitting on screen). Jus the top line: User 51%, System 23%, IOW 0%, IRQ 0%