The simplest solution is to scale to one worker per node initially if you're doing anything compute intensive... once you've done that, and/or you need better performance for any number of reasons including cost, then you can do more. Now, I'm not sure I would have gotten to 4k nodes before I started to re-evaluate parallelism or better scaling options, but the initial implementation is absolutely fine.