This is really intersting.
> In particular, model parallelism is efficient when the amount of computation per neuron activity is high (because the neuron activity is the unit being communicated), while data parallelism is efficient when the amount of computation per weight is high (because the weight is the unit being communicated).
In this case, for fully connected layers, the amount of computation per neuron is high because it is (fully) connected to and from every other neuron in the pervious and following layer, and therefore we want more GPU's to work on those in parallel, so it isn't a time bottleneck?
What does it mean then for to have a high "computation per weight".
>• Convolutional layers cumulatively contain about 90-95% of the computation, about 5% of the parameters, and have large representations.
• Fully-connected layers contain about 5-10% of the computation, about 95% of the parameters, and have small representations.
Wouldn't fc's require more computation since each parameter has more incoming and outgoing values?
I don't doubt this is all accurate and I'm just missing some obvious intuition, but for the sake of learning feel free to explain it if you wish!