What Every User Should Know About Mixed Precision Training in PyTorch (2022)
pytorch.org
pytorch.org
https://pytorch.org/tutorials/intermediate/tensorboard_profi...
Some parts of a module may not work well in lower precision and need to be in higher precision. If you ever create new tensors within the forward pass, you need to manually adjust your code to automatically use the right datatype. You still need to figure out which low-precision dtype you want to use in your model (and perhaps use different ones in different parts of your model). Etc.
Also, I've encountered strange performance regression issues with the newest Docker releases of Tensorflow, with 10x slow-downs compared to previous minor releases. And the docker version was always slower than the local version. Something something Nvidia & CUDA I guess. I had not performance differences with PyTorch when using docker.
It should be said that Tensorflow was generally 10 to 20% faster for similar models. But that could be down to my ineptitude.
As a researcher, Pytorch was also much easier to tinker with, which is perhaps a factor that explains why it rapidly gained popularity in academia.
I switched to Pytorch after I encountered this bug in a very normal use case back in v1.13 https://colab.research.google.com/drive/1D-kgD7NiRXTNTNwVr18...
I've never encountered such a bug in Pytorch in the last 4-5 years.
Zooming in to 125% the effect goes away, and the font-weight at the top and bottom appear equally thick.