That's the smallest of 4 experiments. It goes upto 140 million cells, where MI300X retains similar performance advantage of around 10% over Nvidia's H100.
I've never seen a non-reactive incompressible flow simulation get substantial speedup on GPUs. There are well understood fundamental reasons why this is the case.
Now granted, the flops to byte ratio for this program might be better than an avg fluid simulator. Also, our performance tanked when we moved to multi-node system. But I am aware of underlying reasons behind the scalibility issues and they don't feel like problems that can't be overcome.
Especially if you do the comparison on equivalent cost basis, i.e. "what is the walltime difference if I run on a $60k all-CPU cluster versus a $60k GPU cluster". Or in terms of cloud compute cost / HPC allocation spend.