Using tokens to evaluate models is an outdated approach. Cost per task is what matters. Not all tokens are created equal
All else being equal passing triple the amount of tokens through a model to solve the same problem makes it slower.
Doesn’t mean this model is bad, and it has to be considered how cheap it is, but it’s a factor.