Edit; many of these speeds are low precision which is less useful outside of deep learning, but the higher precision matmul ops in the tensor cores are still very fast and very useful for wide variety of tasks.
The FP64 matrix-multiplication is only 60 TFlops, no where near the advertized 1000 TFlops. TF32 matrix-multiplication is a poorly named 16-bit operation.
I'm on Turing architecture so I've never used TF32. I've only used FP32 and FP16 but FP32 isn't supported by these tensor cores.
Given that it's 32-bit in memory (so all your data structures are 32-bit) and also that in my experience using it is very transparent (I haven't run into any numerical issues compared to full FP32), I think calling it a 32-bit format is a reasonable compromise.
Addition is done in 10-bit mantissa. So maybe TF19 might be the better name, since its a 19-bit format (slightly more than 16-bit BFloats).
Really, its a BFloat with a 10-bit mantissa instead of a 7-bit mantissa. 10-bit mantissa matches FP16, while the 8-bit exponent matches FP32.
So TF19 probably would have been the best name, but NVidia like marketing so they call it TF32 instead.
Yes, the system will read/write the 32-bit value to RAM. But if there's only 10-bits of mantissa in the circuits, you're only going to get 10-bits of precision (best case). The 10-bit mantissa makes sense because these systems have FP16 circuits (1 + 5-bit exponent + 10-bit mantissa) and BFloat16 circuits (1 sign + 8-bit exponent + 7-bit mantissa). So the 8-bit exponent circuit + 10-bit mantissa circuit exists physically on those NVidia cores.
-------
But the 'Tensor Cores' do not support 32-bit (aka: 23-bit mantissa) or higher.
The effective mantissa is like FP16 but it's padded out to be the same size as FP32.
In other words, there's 1 sign bit, 8 exponent bits, 10 mantissa bits that are USED, and 13 mantissa bits that are IGNORED.
1 + 8 + 10 + 13 = 32
The 13 ignored mantissa bits are part of the memory image: they pad the number out to 32-bit alignment.
Only the ppas from graphics-drivers work properly
My experience on windows is much more automatic and it never breaks anything. But I'd rather pay the price (installing on Linux) to avoid windows at all costs
I highly recommend sticking with one technique or the other; never intermix them.
I wish people would stop talking rubbish about NVIDIA's Linux support.