Hi, I’m the author, The Opus 5 substitution was a validation test, not the primary measurement. At that sample size the accuracy difference was -3.8 ± 6.3 points, so it did not clear the pre-registered 99% threshold. Interestingly, output tokens moved much more (-23%), which is why token usage is tracked as a secondary signal.
The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.
The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.