If it can’t even tell apart Opus 5 and 5.5 (according to the readme) then it’s not useful
The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.
The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.