This is actually not about abusively using the computing power (I believe you can use it on a 8-GPU DGX or your DIY workstation), but how to use the same number of GPUs to train a larger model with a larger input size. Pipeline model parallelism seems to be a promising method to increase accuracy by enlarge the model/data size, IIUC.