I really don't get these companies posting disingenuous benchmarks. Every time, they pick and choose who to compare against. Not comparing to the latest 5.3-codex is absurd when it's been out a couple of weeks now. Who are they trying to kid?
People who do not know how reproducible research works.
Any benchmark that is presented by AI labs must be reproduced reliably by someone else independent of that AI lab presenting these results.
Otherwise, not only it is biased, these numbers can be just made up for marketing purposes.
SWE bench for example creates a predictions file and evaluates the results in the harness. Without Codex 5.3 being in the API, it can't.