The number is for FP64 double precision terraflops. For vector operations, not matrix operations as well.
Double precision is important for a lot of simulation codes because of numerical issues. For Deep learning applications it's basically unused, single precision or half precision formats dominate training. The same numerical issues don't show up, so this is a lot more efficient.
Additionally a lot of the 'AI' flops are achieved through matrix cores, fitting Deep learning the workload well. This brings an A100 from only 17 to ~200 tflops.
> But can a 10-node-iphone-cluster really match an A100 for pure flops?!
The iphone number is probably fp32 if not lower precision. However the combined SoCs of 10 iphones have tripple the transistor count and a larger chip area among them. So the comparison is not totally off.