From my experience new models are slower and use more tokens even on questions which gpt 4 answered correctly. It is mostly because newer models tend to be more verbose (even with prompt requesting short answers).
Unless somebody improved on the underlying transformer architecture... Surely AI is smart enough to do it by now