9 → 20 → 36 → 50
These are super arbitrary numbers of benchmarks to run (9, 11, 16, 14). I see a few explanations for this:
> Claude dropped "pathological trials" for you before publishing. This would make your results quite dishonest.
> You chose different sized task groups (almost certainly not the case).
> You excluded timed out trials, or something (also would be dishonest).
Hopefully there's a better explanation here!