It's not because a model performs better in some applications (often by fine-tuning to get better scores at specific tests) that it is better across the board or that we have to believe the company releasing the model with a high number 3 > 2 so that it is commonly accepted as better.
Pushing the reasonnning further: f you need an Opus level performance then not accepting GPT 3 isn't a smell.