Why Does No One Use Advanced Hyperparameter Tuning?
towardsdatascience.com
towardsdatascience.com
A second way to exploit parallelism is via hyperparameter search, and in particular when leveraging early-stopping based approaches like Hyperband/ASHA. Speedups in this setting tend be fairly robust to the the choice of network. As you mentioned, in such settings you can do meaningful amounts of work independently, and moreover can leverage asynchrony to further avoid bottlenecks.
Disclaimer: I am one of the co-founders of Determined, and also one of the developers of Hyperband/Asha.
You can do either data or model parallelism - TensorFlow 2.x has nice utilities for data parallel. DeepSpeed from Microsoft helps with model parallel training. Slow parameter sync between nodes is a bottleneck.