So your process is limited to the resources of a node right? And coordinating data between jobs is via a shared file system or network messaging between nodes?
I'm not sure how big the HMC jobs get on these new machines---it depends on the size of the lattice (which gets optimized for physics but also algorithmic speed / sitting in a good spot for the efficiency of the machine).