But the end user doesn't care about flops for closed models. All they care about is how much it ends up costing them.
But if we're actually talking benchmarking intelligence -- as evidenced by "Artificial Analysis Intelligence Index" and the rest of the reporting -- then comparing systems with different amounts of recurrence on output tokens is flawed.