It's running a very small, non-reasoning model at the moment. But more generally, almost all LLMs argue on the hardware/model they are/are on.
There’s no difference in the inference implementation, parameter count, or speed.
But yeah, there are a lot of factors, so it's hard to answer, and tokens/s isn't the right question.