You mean, comparing the performance after training a model?
These benchmarks aim to highlight the performance differences in terms of speed/memory usage across frameworks and machine configurations.
There is also the practical hurdle that training imagenet models to maximum accuracy takes 1 week+.
Sorry, if I am asking stupid questions.
But in the end, if you are using the same model, the same solver and the same RNG, yes the output of all frameworks should be the same. In practice this also mostly holds true, since the stochastic processes involved are geared towards finding a good local minimum, which is the same given a model and a dataset.
[1]: https://en.wikipedia.org/wiki/Stochastic_gradient_descent