> The A100 draws 131W peak during the training run, the M1MAX GPU draws a maximum of 15.4 W. The A100 runs an iteration at 7ms vs 104ms (with SHARK) and 196ms (with TF-Metal)
> The A100 draws 131W peak during the training run, the M1MAX GPU draws a maximum of 15.4 W. The A100 runs an iteration at 7ms vs 104ms (with SHARK) and 196ms (with TF-Metal)
Honestly though I'm more concerned about the other aspects of apple's software: "The OSX Window Server crashes when all GPUs are used to the maximum"
The good news (for us) is that this is all effectively -O1 today; there's still potential for 2-8x speedups over these numbers and that's not even factoring in the inaccessible HW features that maybe one day Apple will expose :crossed-fingers: :)
// IREE dev
It is, but they’re not using any of that. Only the M1 GPU at this time.
//part of nod.ai / SHARK team.
https://timdettmers.com/2020/09/07/which-gpu-for-deep-learni...
It seems that their consumer GPUs also have a gap