It's not really fair to compare a discrete GPU to a mobile GPU, I only provided this as a comparison for someone who maybe has one of these at home. And btw, you are talking about TF32 performance not FP32. TF32 actually uses 16 bits. A100's FP32 performance is actually lower than 3090, it's 19.5 TFLOPS:
https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
FP16 performance is also relevant as a lot of people now train in FP16. The default for pytorch/TF is still FP32.