Indirectly, using recent 50k tpuv5e run info [1], I'd guess it will give 50-70% of H100 in large-scale LLM jobs: MLCommons show that H100 give a bit more than 500 TFLOPs in FP8, and v5e gives 100 TOPs in INT8. v5p has 2.3 times more theoretical OPs, 3 times more interchip bandwidth and 3.3 times more memory bandwidth, so assuming you can extrapolate and given that bandwidth is usually a bottleneck, ~60% seems plausible.
[1] https://cloud.google.com/blog/products/compute/the-worlds-la...
[2] https://github.com/mlcommons/training_results_v3.1