Right, but the tensor cores should be about 10x faster on the compute side, and about 2x the memory bandwidth. GEMM is usually constrained by compute, which is why the tensor cores exist.
These benchmarks are for training, so the expectation is that they are running them in fp16 all the way through. Also, tensor cores can accumulate in fp32 registers with a slight hit to performance.